跳到正文
原文
Reddit · r/LocalLLaMA· /u/Gold-Bat-3225·· 1 天前AI 评分26

InferBench 基准测试:MiMo V2.6 Pro 接近 GPT-6 Astra,开放权重模型表现超预期

MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants

AI 导读

新发布的 InferBench 基准用 20 个场景、2.8k 条对话测试 12 款前沿大语言模型从指令中推断用户优先级的能力;当用户优先级被明确给出时,模型选择最佳选项的准确率为 89%,部分优先级未给出时降至 60%。

来源:Reddit · r/LocalLLaMA · reddit.com