跳到正文
原文
Reddit · r/MachineLearning· /u/heyitsdannyle·· 2 天前AI 评分27

SWE-Race:基于 188 个真实并发缺陷的编程 Agent 评测基准,三款模型结果出炉

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

AI 导读

SWE-Race 是一个收录约 100 个 Python 项目合并 PR 中真实并发缺陷(竞态、死锁、取消问题)的编程 Agent 评测基准,含 188 个任务,每个任务在无网络容器中由项目自带测试打分。

来源:Reddit · r/MachineLearning · reddit.com