关于情感分析无人真正达成一致:人类、专用工具与大语言模型在社交媒体文本上的困境
Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
Himarsha R. Jayanetti · Sivakanesan Dhanushkanda · Shuai Hao · Michael L. Nelson · Michele C. Weigle
中文摘要
社交媒体蕴含大量实时公众情感信息,但广泛使用的情感分析工具往往在未充分理解其局限性的情况下被直接部署。本研究在 100 条推文上,评估了三种专用情感分析工具(TextBlob、VADER、Twitter-roBERTa-base)和三种大语言模型(Qwen3-32B、GPT-OSS-120B、Llama-4-Maverick-17B)与六位人工标注者之间的一致性。研究采用两类统计指标:用于两两比较的 Cohen's kappa,以及用于多标注者比较的 Fleiss' kappa。结果显示,即使在人工标注者之间,一致性也仅为一般水平(fair agreement),凸显情感分析的主观性。无论是人工标注还是自动化工具,在二分类(负面 vs. 非负面、正面 vs. 非正面)下的一致性均高于三分类。Twitter-roBERTa-base 与人工标注的对齐程度最高,优于各类专用情感分析工具和 LLMs,在区分负面 vs. 非负面情感时尤为突出。各 LLM 之间表现出较强的一致性,与人类标注达成中等至较强的一致性,且在正面 vs. 非正面分类上表现更佳。研究指出,面向特定领域的微调对可靠的社交媒体情感分析至关重要,而以人为中心的评估仍是建立高质量金标准标签的核心环节。
关键要点
- 01问题:主流情感分析工具常被直接套用于社交媒体,但其相对人类标注者的可靠性尚不清晰,且情感本身具有主观性
- 02方法:在 100 条推文上,使用 Cohen's kappa 与 Fleiss' kappa 评估 3 个专用工具与 3 个 LLM 相对 6 位人工标注者的一致性
- 03结果:人工标注者之间一致性仅为 fair;Twitter-roBERTa-base 与人类对齐最佳;LLM 间彼此一致性较高,与人类达 moderate 到 good
- 04结论:二分类(正/负 vs. 非)比三分类更可靠;领域专用微调对社交媒体场景不可或缺,人需对齐仍是金标准的关键
解读
尚无解读。
原始英文摘要
arXiv:2610.10318v1 Announce Type: new Abstract: Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.
同方向论文 · cs.CL
查看全部 →EngramEdit:通过条件记忆实现大语言模型的解耦知识更新
2610.10533Prompt 应该做更多:检索指令对嵌入模型(embedding models)的影响
2610.10508PHRBench:面向大语言模型幻觉后推理(后幻觉推理)的行为评测
2610.10455CoTrace:通过 Harness-Model 协同演化训练终端 Agent 的数据配方
2610.10426使用大语言模型实现爱沙尼亚语文档级文本简化
2610.10378面向任务进度的行动学习:从紧凑教师监督中蒸馏小型智能体
2610.10332