Loading patterns…
SWE-bench Pro(SWE-Pro)
Software engineering tasks long enough to take a professional hours or days, drawn from repositories chosen so that memorising the answer is not an option.
In 30 seconds
- What
- Sets long-horizon software engineering tasks across 41 repositories chosen so that remembering the accepted patch is not a route to a good score.
- When to use
- SWE-bench numbers have saturated for the systems you are comparing, or contamination is the specific question being asked.
- Watch out
- It moves the contamination horizon rather than removing it. This is still a fixed published set, and it will date the way its predecessor did.
Ask the AI expert about this pattern
Opens the assistant with your question prefilled. You review it before sending.
SWE-bench Pro: Overview
Software engineering tasks long enough to take a professional hours or days, drawn from repositories chosen so that memorising the answer is not an option.
- 1,865 problems across 41 actively maintained repositories
- Business, B2B and developer-tool codebases, not only popular OSS
- Patches typically span multiple files and substantial rewrites
- Held-out and commercially licensed subsets to resist contamination
- Human-verified problem statements and interfaces
- Built as the answer to SWE-bench saturation and leakage
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.
Also known as: Long-horizon SWE benchmark, Enterprise coding benchmark
References
The papers, specifications, and repositories this pattern is based on.