Dans l'actualité
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
arXiv cs.AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- FriendBench benchmark tests whether AI models can infer if two people know each other from 20-second conversation clips across text, audio, and video.
- Pourquoi ça compte
- Matters for engineers building social AI, emotion recognition systems, and multimodal models that need to understand human relationships from behavioral cues.
- Vigilance
- Top models match human accuracy but use different reasoning: models bias toward "stranger" while humans stay balanced. Visual behavior helps humans more than models.
Écouter ce résumé
- language model
- prompt
- inference
- benchmark
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.