Loading patterns…
METR RE-Bench
Benchmark for measuring performance of frontier model agents on ML research engineering tasks, comparing against human expert capabilities.
In 30 seconds
- What
- Measures frontier model agent performance on ML research engineering tasks against human expert benchmarks across completion rate, code quality, and research insights.
- When to use
- Assessing whether frontier models can handle real research engineering work or comparing capability gains across model versions.
- Watch out
- Task selection bias heavily influences results; tasks too narrow or too similar to training data inflate agent scores relative to real-world research variety.
Ask the AI expert about this pattern
Opens the assistant with your question prefilled. You review it before sending.
METR RE-Bench: Overview
Benchmark for measuring performance of frontier model agents on ML research engineering tasks, comparing against human expert capabilities.
- ML research engineering task evaluation
- Human vs. agent performance comparison
- Frontier model capability assessment
- Research productivity measurement
- Expert-level task benchmarking
- Comprehensive task diversity
Get the Agent Evals field guide
All 25 agent evaluation methods condensed into one guide: which benchmark measures what, when a public score misleads you, and how to build evals out of your own failures. The link arrives with your confirmation, alongside the weekly Agent Architect.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.
References
The papers, specifications, and repositories this pattern is based on.
- METR RE-Bench: Evaluating R&D Capabilities of LLMs (2024)
- Frontier AI R&D Capabilities Assessment - METR
From the engineer behind this catalog
Find out what your evals miss
Measuring an agent is harder than shipping one, and most suites stay green while production drifts. Have your evaluation setup reviewed end to end: what you measure today, what you cannot see yet, and the regressions your current suite would let through.
€750 instead of €1,500, one week, written report and walkthrough call, until 30 September