EASE:用于规避 AI 生成文本检测器的熵自适应分布整形
EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors
Jicheng Zhou · Kahim Wong · Jialong Wang · Jiantao Zhou
中文摘要
AI 生成文本(AIGT)检测对源大语言模型(LLM)的解码选择较为敏感。研究观察到,扰动下一词元(next-token)logits 或调整采样温度会降低检测性能,清晰暴露出检测器在面对解码阶段的分布变化时的脆弱性。基于此,提出 EASE(Entropy-Adaptive Distribution Shaping for Evasion),一种无需训练且与检测器无关的规避 AIGT 检测框架。EASE 直接依据源 LLM 的下一词元分布计算预测熵,并据此自适应调整 logit 扰动与采样温度,既不依赖检测器反馈,也不需要模型微调。在三种源 LLM 和多种检测器上的实验表明,该方法能持续降低检测性能,且文本质量几乎不退化,推理开销可忽略不计。
关键要点
- 01问题:AIGT 检测器对解码阶段的分布变化敏感,logit 扰动和温度调整即可削弱其性能,暴露出解码时脆弱性。
- 02方法:EASE 是一种无需训练、与检测器无关的框架,利用源 LLM 预测熵自适应调整 logit 扰动和采样温度,不依赖检测器反馈或模型微调。
- 03结果:在三种源 LLM 和多种检测器上,持续降低检测性能,文本质量退化可忽略,推理开销可忽略。
- 04局限:依赖源 LLM 下一词元分布的预测熵信号,对不提供此类分布信息的模型不适用。
解读
尚无解读。
原始英文摘要
arXiv:2610.09976v1 Announce Type: new Abstract: AI-generated text (AIGT) detection can be sensitive to the decoding choices of the source large language model (LLM). We observe that perturbing next-token logits or adjusting sampling temperature can reduce detection performance, providing a clear signal of detector vulnerability to decoding-time distribution changes. Building on this observation, we propose EASE (Entropy-Adaptive Distribution Shaping for Evasion), a training-free and detector-agnostic framework for evading AIGT detectors. EASE computes predictive entropy directly from the source LLM's next-token distribution and uses it to adapt both logit perturbation and sampling temperature, without detector feedback or model fine-tuning. Experiments across three source LLMs and multiple detectors demonstrate consistent reductions in detection performance, with negligible degradation in text quality and negligible inference overhead.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332