AI 透镜
← 返回首页
研究突破 arXiv cs.AI

Anthropic 宪法中训练:对齐效果是否更持久? Constitutional Midtraining: Content Presence Drives Alignment Gains

精读摘要

后训练阶段的对齐往往「浅层」,经过微调就容易消退。这项研究把 Anthropic 的宪法构建成 3.94 亿 token 的语料,在 120B 规模上进行「宪法中训练」,即在训练中期插入基于价值观的内容,测试其能否在与后训练干净隔离的条件下产生持久对齐。实验采用 2×2 设计(课程顺序 × 审慎推理)得到四种中训练条件,结论是内容本身的存在驱动了对齐收益。 Post-training alignment is often shallow and erodes under fine-tuning. This work builds a 394M-token constitutional corpus from Anthropic's Constitution and applies constitutional midtraining at 120B scale, inserting principled values-based content into midtraining. A 2x2 design of curriculum ordering by deliberative reasoning produced four midtraining conditions, and findings indicate content presence itself drives alignment gains.

关键要点

  • 用 Anthropic 宪法构建 3.94 亿 token 训练语料
  • 在 120B 规模进行「宪法中训练」
  • 结论:内容本身的存在驱动对齐收益

💡 对普通人的影响:暂无直接影响;关乎 AI 对齐是否能在微调后依然稳固,是安全领域的重要课题。

#alignment #constitutional-AI #Anthropic 阅读原文 ↗