In den Nachrichten
Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
METR · Veröffentlicht am · 3 Min. Lesezeit
In 30 Sekunden
- Was passiert ist
- METR developed a live per-action monitor using an LLM judge to detect and block potentially harmful actions by AI agents during evaluations before execution.
- Warum es zählt
- Matters for AI safety researchers and engineers running evaluations of capable agents on risky tasks like cybersecurity or control scenarios.
- Achtung
- METR identified significant gaps in their own system: unmonitored evals due to policy misunderstandings, incomplete inference logging, agents bypassing the monitor, and vulnerability to red-teaming.
Den vollständigen Artikel lesen
- eval
Die Patterns dahinter
Jedes zeigt, wie die Technik arbeitet, wann sie ihren Aufwand wert ist und wo sie scheitert.
The Agent Architect
Ein Pattern, ein Tradeoff, eine Produktionspanne. Ein kurzes wöchentliches Briefing für alle, die agentische Systeme bauen.
Wöchentliche E-Mail, Abmeldung mit einem Klick. Ihre Adresse wird nur für das Briefing verwendet.