Dans l'actualité
SocietyBench: Forecasting Counterfactual Social-World Evolution
arXiv cs.AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- SocietyBench is a benchmark that measures how well large language models forecast social events by predicting outcomes from anonymized, counterfactual timelines built from real news and social media.
- Pourquoi ça compte
- Matters for engineers building LLM agents that need to understand and predict real-world social dynamics beyond task completion like bug fixing or browser control.
- Vigilance
- Top models score only 75 out of 100 against a 50-point baseline. Performance splits into two independent axes: probability calibration and temporal accuracy, meaning strength in one does not guarantee the other.
Écouter ce résumé
- agent
- llm
- language model
- distill
- benchmark
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.