通往同一答案的漫漫长路:大语言模型在递增推理预算下的认知偏差
The Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models
Obada Kraishan
中文摘要
推理模型在推理时分配额外算力,并将答案呈现为深思熟虑的产物。若这种深思真如人类认知的双过程理论所暗示那样运作,更长的思考理应削弱由快速直觉判断所产生的经典决策偏差。基于一项已建立的基准,本文使用 30 个情境、覆盖锚定(anchoring)、框架效应(framing)、损失厌恶(loss aversion)、承诺升级(escalation of commitment)、可得性(availability)、确认偏差(confirmation)六类偏差,在四个模型家族上开展剂量–反应研究:每个推理模型与一个匹配的同家族非推理模型配对,在 0、1024、4096、8194 token 的思维上限下请求作答,共计 12,350 次 API 调用。由于请求上限并不等同于实际深思量,本文以每次调用实际消耗的推理 token 作为剂量。结果:第一,推理模型并不比其同家族非推理模型更少偏差;在所有家族中点估计都偏向另一侧,但在题目层面的合并对比不具统计可靠性(Δ = +0.031, t(29) = 1.45, p = .157)。第二,偏差幅度并未随实际深思量增加而可靠下降:没有任何斜率显著为负,凡有变化,反而是有符号分数进一步偏离人类方向。第三,锚定是唯一朝向人类方向的偏差(d = 1.89);其余五类中,有四类在全部七个模型上都偏向相反方向;在每类偏差仅五题的情况下,框架效应方向反转可靠,承诺升级、确认偏差与损失厌恶呈方向性倾向,而可得性偏差则近乎缺失。一条要求在作答前重述锚定值的简短指令,在全部五道锚定题上压低了锚定效应,而再多的思考都做不到这一点,但效果未达显著性(p = .057)。结果表明,不应将测试时推理视为理性的保证,部署模型时仍需逐项审计其偏差。
关键要点
- 01问题:测试时推理(token 级思考预算)被默认等同于更理性的判断,本文质疑这一假设,考察其在六类经典认知偏差上的表现。
- 02方法:在四个模型家族上做剂量–反应研究,每个推理模型配同家族非推理对照,跨 0/1024/4096/8192 token 思维上限、12,350 次 API 调用,以实际消耗推理 token 为剂量。
- 03结果:推理模型并不比同家族非推理模型更少偏差(Δ = +0.031, p = .157);偏差幅度不随深思量增加而下降;锚定是唯一朝人类方向的偏差(d = 1.89),其余四类在七个模型上方向相反。
- 04对照:仅要求在作答前重述锚定值的简短指令,即可在全部五道锚定题上降低锚定效应,远超单纯增加推理 token 的效果,但未达显著性(p = .057)。
- 05局限:每类偏差仅五题,样本较细;结果来自 API 调用的请求–实际剂量差异,通用性有待跨更多模型与情境验证。
解读
尚无解读。
原始英文摘要
arXiv:2610.10049v1 Announce Type: new Abstract: Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking should weaken the classic decision biases that fast, intuitive judgment produces. Using 30 vignettes covering six biases (anchoring, framing, loss aversion, escalation of commitment, availability, confirmation) from an established benchmark, we run a dose-response study across four model families, pairing each reasoning model with a matched non-reasoning sibling and requesting thinking ceilings of 0, 1,024, 4,096, and 8,192 tokens, for 12,350 API calls. Because a requested ceiling is not the same as realized deliberation, we use the reasoning tokens each call consumed as the dose. First, reasoning models are not less biased than their siblings; the point estimate leans the other way in every family, but the item-level pooled contrast is not reliable (Delta = +0.031, t(29) = 1.45, p = .157). Second, bias magnitude does not reliably fall as realized deliberation grows: no slope is significantly negative, and where anything moves it is the signed score drifting further from the human direction. Third, anchoring is the only bias in the human direction (d = 1.89). Four of the other five lean the opposite way in all seven models; with five items per bias, that reversal is reliable for framing and directional for escalation of commitment, confirmation, and loss aversion, while availability is absent. A one-line instruction to restate the anchor before answering lowered anchoring on all five anchoring items, which no amount of additional thinking did, although the effect does not reach significance (p = .057). The results argue against treating test-time reasoning as a rationality guarantee and for auditing deployed models bias by bias.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332