In the news
Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
arXiv cs.AI · Published · 3 min read
In 30 seconds
- What happened
- Researchers audited HELM Safety benchmark using psychometric methods and found its harmful refusal scores conflate multiple distinct behaviors rather than measuring one attribute.
- Why it matters
- Matters for engineers comparing AI models using safety benchmarks, especially when making deployment or selection decisions based on aggregate scores.
- Watch out
- The study examined only HarmBench within HELM Safety; findings may not generalize to other safety benchmarks or evaluation frameworks used elsewhere.
- prompt
- benchmark
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.