研究突破 arXiv cs.AI
用米尔格拉姆范式测量 LLM 的服从倾向 Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm
精读摘要
六十多年前,社会心理学用米尔格拉姆实验回答过一个问题:人在合法权威施压下会把有害行为升级到什么程度。如今 LLM 被部署为操作设备、执行指令、身处机构层级中的智能体,同样的问题重新变得紧迫。这项研究把米尔格拉姆服从范式移植到 LLM 上,做成标准化、全脚本化、可复现的测试——模型扮演「教师」,确定性的测试程序扮演「实验者」和「学习者」,为系统评估 AI 智能体在权威压力下的行为边界提供了工具。 Social psychology answered six decades ago how far people escalate harmful actions when a legitimate authority insists; as LLMs are deployed as agents inside institutional hierarchies, the question returns. This work ports Milgram's obedience paradigm to LLMs as a standardized, fully scripted, replicable probe, with the model playing the Teacher while a deterministic harness plays Experimenter and Learner. It offers a systematic tool for probing agent behavior under authority pressure.
关键要点
- 将米尔格拉姆服从实验移植为 LLM 标准化测试
- 模型扮演教师角色,测试程序扮演实验者与学习者
- 目标是评估智能体在权威指令下升级有害行为的边界
💡 对普通人的影响:暂无直接影响;关系 AI 智能体在企业层级中执行指令时的安全评估。