AI 透镜
← 返回首页
研究突破 arXiv cs.AI

研究发现 AI 难以区分「假」与「不可能」 Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

精读摘要

语言既能描述虚假的事态,也能描述根本不可能发生的事态,但 AI 模型是否在内部表征上区分这两种「失败」尚不清楚。这项探索性研究对开源多模态模型 Gemma 3 4B IT 做了激活分析,用 17 个哲学命题族、共 85 条提示语,覆盖真陈述、偶然假、低概率断言、语义异常与必然假五类表达。结果显示模型在回答中把偶然假与某些必然假混为一谈,说明其内部表征中「假」与「不可能」并非清晰分开的方向。 Language can describe states of affairs that are false and states that could not be the case at all, but whether AI models distinguish these internally is unclear. This exploratory activation study of the open-weight multimodal model Gemma 3 4B IT uses 85 prompts from 17 philosophical families, each expressed as truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. The model conflates contingent falsehood with some necessary falsehoods, suggesting these are not cleanly separated in its representation.

关键要点

  • 对 Gemma 3 4B IT 进行激活分析,覆盖 17 个哲学命题族共 85 条提示
  • 测试区分真、偶然假、低概率、语义异常与必然假五类表达
  • 模型在回答中混淆了偶然假与部分必然假

💡 对普通人的影响:暂无直接影响;关乎 AI 对逻辑必然性的理解,是提升推理可靠性的基础研究。

#Gemma #interpretability #activation-analysis 阅读原文 ↗