In the news
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
arXiv cs.AI · Published · 3 min read
In 30 seconds
- What happened
- Researchers found that GPT models reduce explicit harmful language while shifting gender discrimination into subtler forms that toxicity classifiers miss.
- Why it matters
- Matters for engineers building safety evaluations, deploying language models, or relying on automated harm metrics to verify model improvements.
- Watch out
- Standard toxicity scoring tools may show improvement while representational harms actually grow; surface-level metrics alone cannot validate safety progress.
- language model
- eval
- gpt
- phi
The patterns behind this
- Constitutional Classifiers
- Generative UI (Agent-Rendered Interfaces)
- Agentic Context Engineering (Evolving Playbook)
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.