正在加载模式…
SWE-bench Pro(SWE-Pro)
Software engineering tasks long enough to take a professional hours or days, drawn from repositories chosen so that memorising the answer is not an option.
30秒速览
- 是什么
- Sets long-horizon software engineering tasks across 41 repositories chosen so that remembering the accepted patch is not a route to a good score.
- 何时使用
- SWE-bench numbers have saturated for the systems you are comparing, or contamination is the specific question being asked.
- 注意
- It moves the contamination horizon rather than removing it. This is still a fixed published set, and it will date the way its predecessor did.
向AI专家咨询此模式
打开助手并预填您的问题,发送前可先确认。
SWE-bench Pro: 概览
Software engineering tasks long enough to take a professional hours or days, drawn from repositories chosen so that memorising the answer is not an option.
- 1,865 problems across 41 actively maintained repositories
- Business, B2B and developer-tool codebases, not only popular OSS
- Patches typically span multiple files and substantial rewrites
- Held-out and commercially licensed subsets to resist contamination
- Human-verified problem statements and interfaces
- Built as the answer to SWE-bench saturation and leakage
获取智能体评估实战指南
25 种智能体评估方法浓缩成一份指南:每个基准测量什么、公开分数何时具有误导性、如何用自己的失败构建评估。链接随确认邮件送达,并附每周的 The Agent Architect。
每周一封邮件,一键退订。您的地址仅用于发送简报。
又称: Long-horizon SWE benchmark, Enterprise coding benchmark
参考资料
该模式所依据的论文、规范和代码仓库。
由本目录背后的工程师执行
看看你的评测漏掉了什么
评测智能体比做出来更难,多数测试套件一路全绿,线上却在悄悄漂移。我们从头到尾检查你的评测体系:现在测到了什么、还有什么看不见,以及当前套件会放过哪些回归。
€750(原价 €1,500),一周交付,书面报告加讲解通话,9月30日前