Loading...
METR RE-Bench(RE-Bench)
Benchmark for measuring performance of frontier model agents on ML research engineering tasks, comparing against human expert capabilities.
In 30 seconds
- What
- Measures frontier model agent performance on ML research engineering tasks against human expert benchmarks across completion rate, code quality, and research insights.
- When to use
- Assessing whether frontier models can handle real research engineering work or comparing capability gains across model versions.
- Watch out
- Task selection bias heavily influences results; tasks too narrow or too similar to training data inflate agent scores relative to real-world research variety.
Loading technique guide…