Dans l'actualité
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
arXiv cs.AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- ERUnderstand benchmark evaluates vision-language models on understanding Entity-Relationship diagrams, with 2,960 labeled diagrams and machine-readable schema representations.
- Pourquoi ça compte
- Database engineers and AI researchers building tools to automate schema extraction from diagram images need standardized evaluation metrics.
- Vigilance
- Models struggle with complex constructs: weak entities score 0.28 F1, multivalued attributes 0.14 F1, N-ary relationships 0.07 F1 despite strong common element recovery.
Écouter ce résumé
Extrait de l'article
-->
Computer Science > Artificial Intelligence
arXiv:2607.24707v1 (cs)
[Submitted on 27 Jul 2026]
Title: ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
Authors: Ali Ansari , Yasmin Mohammadi , Farnoush Nili , Parsa Esmaeilkhani , Longin Jan Latecki , Eduard Dragut
View a PDF of the paper titled ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams, by Ali Ansari and 5 other authors
View PDF HTML (experimental)
Abstract: Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated examples spanning diverse domains, notations, complexity levels, and Extended Entity-Relationship (EER) constructs. Each diagram is paired with a standardized machine-readable representation for fine-grained evaluation of schema elements. Evaluating state-of-the-art Vision-Language Models (VLMs), we find that while common ERD elements are recovered reliably (F1 > 0.74), performance drops sharply on weak entities (as low as 0.28 F1), multivalued attributes (0.14 F1), and N-ary r
Extrait de l'original. Lisez l'article complet à la source.
- language model
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.