跳到正文
返回论文列表
cs.AI提交于 已译

无地面真理下的有效性:陈述偏好经济学为语言模型评估提供了什么

Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

Daniel Robert Kling Alexander · Catherine Louise Kling

中文摘要

向大语言模型提出的许多问题没有可对照评分的正确答案:一项政策值多少、用户应选择哪个选项、如何权衡相互竞争的价值。陈述偏好经济学数十年来一直面临这一难题。它通过一套有效性及相关概念框架——内容,种树、构念,标尺与准则有效性、信度、激励相容与结果性——来评判无确定值的回答。本文主张,这一框架是评估语言模型的通用方法,并阐明每个概念在 LLM 评估中的含义。借助 Vossler 等人 2023 年发表的一项已发表的水质陈述偏好经济价值评估调查,本文在六个模型上开展了实证。在该经济学应用中,有效性检验采取经济学理论预测的形式:需求曲线应向下倾斜,支付意愿应随物品范围与收入而变化。这些检验明显区分了模型表现:两个较早的模型在 75,000 美元家庭收入水平上未能通过最基本的检验;两个最新的模型通过了所有可评分的理论有效性检验,但在收敛有效性上出现分歧。通过有效性检验仅能说明模型回答具有内部连贯,而非其正确。

关键要点

  1. 01问题:大语言模型面临的许多提问无地面真理,传统基于正确答案的打分方法失效
  2. 02方法:将陈述偏好经济学中的内容、构念、准则三种有效性及激励相容等概念引入 LLM 评估
  3. 03方法:以水质支付意愿调查为案例,依据经济学理论预测对六个模型做有效性检验
  4. 04结果:两个旧模型在 $75,000 收入档未通过最基本检验,两个新模型通过所有理论有效性检验但收敛有效性分歧
  5. 05局限:通过有效性检验只证明回答的内部连贯性,不等于回答正确,无法替代正确性评估

解读

尚无解读。

原始英文摘要

arXiv:2610.10506v1 Announce Type: cross Abstract: Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of \$75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model's answers are coherent, not that they are correct.

同方向论文 · cs.AI

查看全部 →