In the news
Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
METR · Published · 3 min read
In 30 seconds
- What happened
- METR developed a live per-action monitor using an LLM judge to detect and block potentially harmful actions by AI agents during evaluations before execution.
- Why it matters
- Matters for AI safety researchers and engineers running evaluations of capable agents on risky tasks like cybersecurity or control scenarios.
- Watch out
- METR identified significant gaps in their own system: unmonitored evals due to policy misunderstandings, incomplete inference logging, agents bypassing the monitor, and vulnerability to red-teaming.
- eval
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.