METR published research on metrics for measuring agent ability.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →METR published analysis on the economics of recursive self-improvement.
METR published research on measuring optimization ability using NanoGPT as an application.
Title alone insufficient for concrete summary.
METR published predeployment evaluation results for GPT-5.6 Sol.
METR published a frontier risk report covering February to March 2026.
METR measured the self-reported impact of early-2026 AI on technical worker productivity.
METR reviews the automated R&D risks section from Anthropic's February 2026 Risk Report.
METR studied AI R&D progress using NanoGPT as evidence.
Fine-tuning experiments investigate chain-of-thought controllability.
METR red-teamed Anthropic's internal agent monitoring systems.
METR reviews Anthropic's sabotage risk report for Claude Opus 4.6.
METR observed results from two CLI game reimplementation runs using Anthropic's Opus 4.6 model.
Researchers analyzed coding agent transcripts to estimate upper bounds on productivity gains from AI agents.
METR measured time horizon using Claude Code and Codex.
METR published early work on evaluating monitorability of AI systems.
METR updated documentation on common elements of frontier AI safety policies as of December 2025.
METR published evaluation results for GPT-5.1-Codex-Max model.
METR reviews Anthropic's Summer 2025 pilot sabotage risk report.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.