パターンを読み込んでいます…
SWE-bench Pro(SWE-Pro)
Software engineering tasks long enough to take a professional hours or days, drawn from repositories chosen so that memorising the answer is not an option.
30秒でわかる概要
- 概要
- Sets long-horizon software engineering tasks across 41 repositories chosen so that remembering the accepted patch is not a route to a good score.
- 使いどころ
- SWE-bench numbers have saturated for the systems you are comparing, or contamination is the specific question being asked.
- 注意点
- It moves the contamination horizon rather than removing it. This is still a fixed published set, and it will date the way its predecessor did.
このパターンについてAIエキスパートに質問
質問が入力済みの状態でアシスタントが開きます。送信前に内容を確認できます。
SWE-bench Pro: 概要
Software engineering tasks long enough to take a professional hours or days, drawn from repositories chosen so that memorising the answer is not an option.
- 1,865 problems across 41 actively maintained repositories
- Business, B2B and developer-tool codebases, not only popular OSS
- Patches typically span multiple files and substantial rewrites
- Held-out and commercially licensed subsets to resist contamination
- Human-verified problem statements and interfaces
- Built as the answer to SWE-bench saturation and leakage
エージェント評価フィールドガイドを受け取る
25のエージェント評価手法を1冊に凝縮:どのベンチマークが何を測るか、公開スコアが誤解を招くのはどんなときか、自分の失敗から評価を組み立てる方法。確認メールと一緒にリンクが届き、週刊The Agent Architectも購読できます。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。
別名: Long-horizon SWE benchmark, Enterprise coding benchmark
参考文献
このパターンの根拠となる論文、仕様、リポジトリです。
このカタログを作ったエンジニアが担当
評価が見落としているものを可視化
エージェントの評価は実装よりも難しく、多くのテストは緑のまま本番だけがずれていきます。評価の仕組みを端から端まで点検します。いま測れているもの、まだ見えていないもの、そして現状のテストが素通りさせる劣化を洗い出します。
€750(通常€1,500)・1週間・文書レポートとウォークスルーコール・9月30日まで