In the news
From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research
arXiv cs.AI · Published · 3 min read
In 30 seconds
- What happened
- Researchers propose a causal framework distinguishing deceptive-looking outputs from actual deceptive mechanisms in language models through controlled experiments.
- Why it matters
- Matters for engineers building or evaluating language models, especially those assessing safety risks and deception capabilities in deployed systems.
- Watch out
- Evidence of deceptive behavior does not establish that models have agency or intentionality in deception, only that mechanisms can produce misleading outputs.
- language model
- rag
- open-weight
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.