PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
Linghao Meng · Feng He · Xuan Yang · Junyuan Mao · Pinze Ren · Deqing Mu · et al.
中文摘要
在大语言模型(LLM)多阶段系统中,幻觉信息可沿推理链向下传递,成为后续推理的上下文。既有针对后幻觉推理(post-hallucination reasoning, PHR)的研究多聚焦最终输出与整体推理动态的统计变化,缺乏对模型如何在单条响应层面解析错误前提的行为级刻画。本文提出 PHRBench,一个跨四个领域、覆盖 18 个大语言模型的受控行为评测基准。该基准在剥离最终答案正确性的前提下,通过幻觉顺从(Hallucination Compliance)、幻觉规避(Hallucination Avoidance)与启发式纠错(Heuristic Correction)三类指标独立刻画每条推理轨迹,并将最终纠正至正确答案的轨迹定义为有效推理轨迹。在 4820 条受控样本上,成功纠错整体仍属少数,且与推理过程中更频繁的信念更新相关。进一步分析表明,被污染提示(prompt)的属性对成功纠错具有显著预测能力,一个轻量分类器即可达到 0.847 的 AUROC。本工作从行为视角刻画了大语言模型解析错误上下文的方式及成功纠错的发生条件。
关键要点
- 01提出 PHRBench,首个跨四个领域、覆盖 18 个大语言模型的后幻觉推理受控行为评测基准
- 02通过幻觉顺从、幻觉规避、启发式纠错三类指标,在 4820 条样本上独立于最终答案正确性刻画推理轨迹行为
- 03发现成功纠错在 LLM 中仍属少数,且与推理过程中更频繁的信念更新相关
- 04幻觉化提示的属性对成功纠错具有强预测信号,轻量分类器 AUROC 达 0.847
- 05研究局限:仅评估 18 个模型且聚焦行为层面,未涉及幻觉产生机理或多阶段系统的传播路径
解读
尚无解读。
原始英文摘要
arXiv:2610.10455v1 Announce Type: new Abstract: Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332关于情感分析无人真正达成一致:人类、专用工具与大语言模型在社交媒体文本上的困境
2610.10318