跳到正文
原文
Reddit · r/artificial· /u/radeon2000·· 4 天前AI 评分67

我让13个AI模型玩医疗咨询游戏:全部诊断正确,但安全表现差异大

I made 13 AI models play the doctor in my medical consultation game. All 195 consults got the diagnosis right; what separated them was safety.

AI 导读

作者让13个AI模型在自建的医疗咨询游戏中进行诊断测试,所有195次咨询都给出了正确诊断(心梗、阑尾炎、肺炎),但安全评估才是真正的区分点。GPT-6 Astra以83%得分和88%红旗识别率领先,GPT-6.1 Sol以80%得分和仅$0.03/次的成本表现亮眼,而Claude Fable 5.1得分75%但成本高达$2.06/次。

来源:Reddit · r/artificial · reddit.com