Reddit · r/MachineLearning· /u/heyitsdannyle·· 2 天前AI 评分27
SWE-Race:基于 188 个真实并发缺陷的编程 Agent 评测基准,三款模型结果出炉
SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
AI 导读
SWE-Race 是一个收录约 100 个 Python 项目合并 PR 中真实并发缺陷(竞态、死锁、取消问题)的编程 Agent 评测基准,含 188 个任务,每个任务在无网络容器中由项目自带测试打分。
来源:Reddit · r/MachineLearning · reddit.com