Dans l'actualité
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
arXiv cs.AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- GPT models reduce toxicity scores but transform rather than eliminate gender discrimination across generations, missing representational harms.
- Pourquoi ça compte
- Critical for engineers building LLM safety systems who use automated toxicity metrics to validate model improvements.
- Vigilance
- Standard toxicity classifiers miss representational harms; declining scores may mask bias pattern shifts rather than genuine harm reduction.
- language model
- eval
- gpt
- phi
Les patterns derrière cette actualité
- Constitutional Classifiers
- Generative UI (Agent-Rendered Interfaces)
- Agentic Context Engineering (Evolving Playbook)
Chacun explique le fonctionnement de la technique, quand elle vaut son coût et où elle casse.
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.