跳到正文
返回论文列表
cs.CL提交于 已译

LLM 说服效果取决于评估视角

LLM Persuasion Is in the Eye of the Evaluation

Kamile Dementaviciute · Julija Vaitonyte · Tijl De Bie

中文摘要

已有研究表明,大语言模型(LLMs)在说服任务上已达到甚至超越人类专家水平。其说服能力在教育、健康传播等场景中具有积极应用潜力,但同样可被用于操纵与误导,因此对其评估已成为日益迫切的需求。然而,当前评估仍缺乏统一标准:不同研究对"说服"的定义不一,许多宽泛结论往往建立在狭窄且局限于特定情境的评估之上。自动化评估为横向对比提供了途径——可在同一批模型上大规模运行,并能覆盖难以或不应对人类测试的某些高风险说服形式。本研究将9种已发表的自动化方法统一到同一框架下,在15个LLM上分别运行,考察这些方法的排名是否一致以及原因。结果显示,方法间一致性较弱(Spearman秩相关系数均值约为0.393——注:原文给出均值 Spearman ρ=0.25)。分析揭示两方面原因:对部分任务拒绝回答、其余任务不拒绝的模型(无论直接还是间接)使一致性下降约四分之一,且拒绝现象集中在操纵任务上;通用能力同样产生影响——多数理性说服(非操纵)方法的排名与模型通用能力一致,多数操纵方法则不然。综合而言,这些发现表明一致性更多取决于方法所设定的任务,而非其评分方式,但鉴于仅有8种方法可纳入分析,该结论仅为初步指示。更广泛地看,说服评分反映的是模型"说服能力"与"说服意愿"的共同作用;单一评分仅在其自身设定下具有参考意义,难以全面反映模型在跨任务中的说服水平。

关键要点

  1. 01问题:现有LLM说服力评估方法定义不一、缺乏统一标准,大量宽泛结论建立在窄场景测试上,方法间可比性差
  2. 02方法:把9种已发表自动化说服评估方法统一到同一框架,在15个LLM上运行,横向对比其排名一致性及成因
  3. 03结果:9种方法排名一致性较弱(Spearman ρ均值0.25),说明单一评分难以反映模型的跨任务说服能力
  4. 04结果:任务拒绝使一致性下降约四分之一,集中在操纵类任务;理性说服方法的排名与通用能力更相关,操纵方法则不然
  5. 05局限:纳入一致性分析的方法仅8种,关于"任务设计主导评分差异"的结论仅为初步,且尚未覆盖所有说服场景

解读

尚无解读。

原始英文摘要

arXiv:2610.10232v1 Announce Type: new Abstract: Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman $\rho = 0.25$). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.

同方向论文 · cs.CL

查看全部 →