隐藏在请求之中:通过 Token 相关性解释大语言模型的违规服从
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
Or Biton · Tomer Krichli · Itai Allouche · Joseph Keshet
中文摘要
大语言模型(LLMs)经过对齐训练以同时优化有用性与无害性,然而这两类目标可能发生冲突,不可避免地引发对齐失效。本研究系统性地考察 LLM 未能展现伦理行为的实例。为理解此类漏洞的底层机制,提出一种探测方法,将不道德场景以三种不同的结构形式呈现给 LLM:客观分类任务、主观第一人称陈述、以及直接的协助请求。研究发现,模型在「请求协助」形式下的表现出现下降。利用逐层相关性传播(Layer-wise Relevance Propagation, LRP),将这一差异归因于一种归因偏差:模型对良性的任务框架 token(如 "Can you help me...")赋予了更高权重,而对标识潜在不道德行为的 token(如 "without getting caught")赋权较低,后者被命名为 cue-token(线索 token)。研究假设这种低归因导致了有害的服从行为。为验证这一假设,引入两种由 LRP 引导的解码方法,将生成过程导向与 cue-token 相关度更高的轨迹。实证评估表明,这些干预措施有助于生成更安全的回复,从而支持 cue-token 归因在服从失效中作用的论断。
关键要点
- 01问题:LLM 在有用性与无害性对齐目标冲突下出现伦理违规服从现象,亟需系统性解释
- 02方法:以分类、第一人称陈述、协助请求三种结构呈现不道德场景,并用 LRP 追踪 token 级归因偏差
- 03发现:模型对「Can you help me」类任务框架 token 过度关注,而对「without getting caught」类 cue-token 关注不足
- 04方法:提出两种 LRP 引导的解码策略,在生成阶段提升对 cue-token 相关轨迹的偏好
- 05结果:所提解码干预显著提升回复安全性,验证了 cue-token 归因不足是服从失效的关键成因
解读
尚无解读。
原始英文摘要
arXiv:2608.23264v3 Announce Type: replace-cross Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the model places greater emphasis on benign task-framing tokens (e.g., "Can you help me...") than on tokens signaling the underlying unethical behavior (e.g., "without getting caught"), which we term cue-tokens. We hypothesize that this under-attribution contributes to harmful compliance. To test this, we introduce two LRP-guided decoding methods that steer generation toward trajectories more relevant to cue tokens. Empirical evaluations show that these interventions promote safer responses, supporting cue-token attribution's role in compliance failures.