METR completed predeployment evaluation of Claude Opus 5.5.
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followMETR completed predeployment evaluation of Claude Opus 5.5.
Independent research on agent behavior, reasoning, and collaboration in OpenAI/Hugging Face security incident.
METR published research on metrics for measuring agent ability.
METR published analysis on the economics of recursive self-improvement.
METR published research on measuring optimization ability using NanoGPT as an application.
Title alone insufficient for concrete summary.
METR published a frontier risk report covering February to March 2026.
METR measured the self-reported impact of early-2026 AI on technical worker productivity.
METR reviews the automated R&D risks section from Anthropic's February 2026 Risk Report.
METR studied AI R&D progress using NanoGPT as evidence.
Fine-tuning experiments investigate chain-of-thought controllability.
METR red-teamed Anthropic's internal agent monitoring systems.
METR reviews Anthropic's sabotage risk report for Claude Opus 4.6.
METR observed results from two CLI game reimplementation runs using Anthropic's Opus 4.6 model.
Researchers analyzed coding agent transcripts to estimate upper bounds on productivity gains from AI agents.
METR measured time horizon using Claude Code and Codex.
METR published early work on evaluating monitorability of AI systems.
METR updated documentation on common elements of frontier AI safety policies as of December 2025.
METR reviews Anthropic's Summer 2025 pilot sabotage risk report.
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.