宁愿退出NLP也不愿再读一篇这样的论文:NLP论文中对立结构(antithesis)的兴起
I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers
Olga Zamaraeva · Adri\'an Gude · Roi Santos-R\'ios · Carlos G\'omez-Rodr\'iguez
中文摘要
无论好坏,大语言模型(LLM)如今已常规用于学术写作。许多人注意到,近期模型在论文中大量堆砌不必要的对立结构,反复强调工作不做什么,这类表达既无助于精确性也无助于表达质量,反而令审稿人厌烦而非留下深刻印象。研究考察了rather than这一结构在2019年ACL论文、2026年ACL风格arXiv论文,以及由GPT模型根据相同标题与摘要生成的论文中的使用情况。2026年的使用率是2019年的七倍,在GPT论文中更高。两名不知来源的标注者认为2019年的用例几乎都不令人厌烦,2026年的约十分之一令人厌烦;两人很少对具体哪一例意见一致,但2026年论文中约有一半各含一例令其各自厌烦的用法。令人厌烦的用法比正当用法更倾向于将所否定的替代项描述得更为有利。在偏好数据与开源奖励模型(reward model)评估中,评分者更偏好这一结构,且「请保持诚实」之类的指令反而会助长其使用。作者推测这是成对偏好( pairwise preferences )后训练的副作用:该训练对单一回复中的自我否定给予奖励,却无法衡量其贯穿全文的代价。
关键要点
- 01问题:LLM生成的论文中存在大量不必要的rather than对立结构,影响表达质量并引起审稿人反感
- 02方法:对比2019年ACL论文、2026年ACL风格arXiv论文与GPT生成论文中rather than的使用率,由盲标注者评判
- 03结果:2026年使用率约为2019年的七倍,2026年论文约一半各含一例令标注者厌烦的用法,烦人用法更倾向贬低被否定的替代项
- 04分析:偏好数据与奖励模型更偏好该结构,「诚实」指令反而促进其使用,推测源于成对偏好后训练无法衡量全文代价
解读
尚无解读。
原始英文摘要
arXiv:2610.10092v1 Announce Type: new Abstract: For better or worse, LLMs are by now used routinely for scientific writing.\footnote{This paper is no exception; we did use AI to assist with writing some of the sections (see Acknowledgments).} Many have noticed that recent models fill papers with unnecessary antithesis, stating over and over what the work does not do, in ways that do not contribute to its precision or quality of expression and annoy reviewers \emph{rather than impressing them}. We study the construction \emph{rather than} in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts. Its rate in 2026 is seven times the 2019 rate, and higher still in the GPT papers. Two annotators, blind to the source, find almost no 2019 use \emph{annoying} and about one in ten 2026 uses; they seldom agree on which, yet about half of 2026 papers contain a use that annoys each of them. \emph{Annoying} uses present the rejected alternative less favorably than legitimate uses. Raters of preference data and open reward models favor the construction, and an instruction to be honest promotes it. We conjecture that it is a side effect of post-training on pairwise preferences, which credit a disavowal in a single response and cannot register its cost across a text.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332