Loading patterns…
SWE-bench Suite
Software engineering benchmark suite including SWE-bench, SWE-bench Verified, and SWE-bench Live. Named comparison models in the example are historical baselines.
In 30 seconds
- What
- Evaluates coding agents against real GitHub issues from popular repositories, with human-validated subsets and monthly-updated live variants to measure problem-solving capability.
- When to use
- Comparing agent performance on authentic software engineering tasks, tracking improvement over time, or validating that agents solve problems without training data contamination.
- Watch out
- Results depend heavily on which repositories and issue types are included; performance on SWE-bench may not transfer to your codebase's patterns or tech stack.
Ask the AI expert about this pattern
Opens the assistant with your question prefilled. You review it before sending.
SWE-bench Suite: Overview
Software engineering benchmark suite including SWE-bench, SWE-bench Verified, and SWE-bench Live. Named comparison models in the example are historical baselines.
- Real GitHub issues from 12+ popular repositories
- Human-validated problem subset (SWE-bench Verified)
- Live benchmark updated monthly (SWE-bench Live)
- Multimodal variant with visual elements
- Contamination-free evaluation
- Industry standard for coding agents
Get the Agent Evals field guide
All 25 agent evaluation methods condensed into one guide: which benchmark measures what, when a public score misleads you, and how to build evals out of your own failures. The link arrives with your confirmation, alongside the weekly Agent Architect.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.
References
The papers, specifications, and repositories this pattern is based on.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (ICLR 2024)arXiv:2310.06770
- Introducing SWE-bench Verified - OpenAI (2024)
- swebench.com
From the engineer behind this catalog
Find out what your evals miss
Measuring an agent is harder than shipping one, and most suites stay green while production drifts. Have your evaluation setup reviewed end to end: what you measure today, what you cannot see yet, and the regressions your current suite would let through.
€750 instead of €1,500, one week, written report and walkthrough call, until 30 September