In den Nachrichten
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
arXiv cs.AI · Veröffentlicht am · 3 Min. Lesezeit
In 30 Sekunden
- Was passiert ist
- FriendBench benchmark tests whether AI models can infer if two people know each other from 20-second conversation clips across text, audio, and video.
- Warum es zählt
- Matters for engineers building social AI, emotion recognition systems, and multimodal models that need to understand human relationships from behavioral cues.
- Achtung
- Top models match human accuracy but use different reasoning: models bias toward "stranger" while humans stay balanced. Visual behavior helps humans more than models.
Diese Zusammenfassung anhören
Den vollständigen Artikel lesen
- language model
- prompt
- inference
- benchmark
The Agent Architect
Ein Pattern, ein Tradeoff, eine Produktionspanne. Ein kurzes wöchentliches Briefing für alle, die agentische Systeme bauen.
Wöchentliche E-Mail, Abmeldung mit einem Klick. Ihre Adresse wird nur für das Briefing verwendet.