跳到正文
原文
Reddit · r/LocalLLaMA· /u/More_Slide5739·· 6 天前AI 评分58

无需VRAM检测本地模型幻觉:1.5B到120B模型测试结果

Detecting hallucinations in local models without eating VRAM: What we learned testing 1.5B to 120B models

AI 导读

作者测试了一种零GPU消耗的幻觉检测方法(Spanda,基于确定性字符串规范化+香农熵),在GSM8K和TriviaQA上对1.5B到120B模型基准测试。结果显示:Qwen 1.5B的AUROC仅0.58(语法太松散),Mistral 7B达0.71(与重型DeBERTa NLI模型相当),Qwen 27B达0.89(精确匹配主导)。

来源:Reddit · r/LocalLLaMA · reddit.com