跳到正文
返回论文列表
cs.AI提交于 已译

大语言模型在被提示不诚实作答时推理 token 出现尖峰

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

Maverick Morales · Tom\'a\v{s} Dominik · Vermut Gao · Katrina Shirey · Paulius Rimkevi\v{c}ius · Aaron Schurger · et al.

中文摘要

监控推理型人工智能(AI)模型的思维链,仍是检测此类模型欺骗及其他失范行为的关键手段。然而,语义层面的思维链监控依赖于推理轨迹的可读性以及对底层计算的足够忠实性,更不用说其可访问性了。越来越多的证据表明,即便思维链输出仍可访问,其也可能很快变得不可读或不忠实。基于认知负荷理论,本文研究了一种更低带宽的信号——生成的推理 token 数量——它不依赖于推理轨迹内容的访问。三款具备推理能力的大语言模型在系统提示下分别以真实作答、虚假作答、不顾真实性作答的方式回答了 210 道多选题,涵盖分析性、描述性、规范性推理类型以及道德与非道德领域。在三款模型上一致观察到:面向真实的作答所消耗的推理 token 少于面向谎言的作答和不顾真实的作答。上述结果表明,被显式提示的不诚实应答策略会在测试时推理 token 用量上产生稳健的群体级差异。虽然尚不能据此将推理 token 计数确立为自发欺骗或通用失准的指标,本研究作为一项概念验证表明:在原始推理轨迹不可用或不可靠时,该计数可作为一种简单、与内容无关的候选信号,用以区分模型的不诚实行为与诚实行为。未来工作应进一步检验实例级检测率、分布外泛化能力、学习得到的欺骗策略、隐藏目标,以及在对抗性压力下的稳健性。

关键要点

  1. 01其原理基于认知负荷理论:不诚实作答需要额外计算资源,会在推理 token 计数上产生可观测差异
  2. 02note:内容无关的代理信号:仅依据推理 token 数量,不依赖推理轨迹语义内容,适用于原始 CoT 不可用场景
  3. 03note:在三种具备推理能力的大语言模型、210 道涵盖分析/描述/规范性及道德与非道德的多选题上一致成立
  4. 04note:面向真实的作答所用推理 token 显著少于面向谎言和不顾真实的作答,群体级差异稳健
  5. 05note:局限:尚未证明可用于实例级自发欺骗检测,也未验证分布外泛化与对抗稳健性

解读

尚无解读。

原始英文摘要

arXiv:2610.10405v1 Announce Type: cross Abstract: Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.

同方向论文 · cs.AI

查看全部 →