Новости
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
arXiv cs.AI · Опубликовано · 3 мин чтения
За 30 секунд
- Что произошло
- ERUnderstand benchmark evaluates vision-language models on understanding Entity-Relationship diagrams, with 2,960 labeled diagrams and machine-readable schema representations.
- Почему это важно
- Database engineers and AI researchers building tools to automate schema extraction from diagram images need standardized evaluation metrics.
- На что обратить внимание
- Models struggle with complex constructs: weak entities score 0.28 F1, multivalued attributes 0.14 F1, N-ary relationships 0.07 F1 despite strong common element recovery.
Послушать это резюме
Из статьи
-->
Computer Science > Artificial Intelligence
arXiv:2607.24707v1 (cs)
[Submitted on 27 Jul 2026]
Title: ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
Authors: Ali Ansari , Yasmin Mohammadi , Farnoush Nili , Parsa Esmaeilkhani , Longin Jan Latecki , Eduard Dragut
View a PDF of the paper titled ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams, by Ali Ansari and 5 other authors
View PDF HTML (experimental)
Abstract: Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted database engineering. We introduce ERUnderstand, the first large-scale benchmark for structured understanding of ER diagrams, comprising 2,960 diagrams collected from curated educational sources, real-world schemas, and synthetically generated examples spanning diverse domains, notations, complexity levels, and Extended Entity-Relationship (EER) constructs. Each diagram is paired with a standardized machine-readable representation for fine-grained evaluation of schema elements. Evaluating state-of-the-art Vision-Language Models (VLMs), we find that while common ERD elements are recovered reliably (F1 > 0.74), performance drops sharply on weak entities (as low as 0.28 F1), multivalued attributes (0.14 F1), and N-ary r
Фрагмент оригинала. Полный текст читайте в первоисточнике.
- language model
The Agent Architect
Один паттерн, один компромисс, одна история сбоя в продакшене. Короткий еженедельный брифинг для тех, кто строит агентные системы.
Одно письмо в неделю, отписка в один клик. Адрес используется только для рассылки брифинга.