Новости
SceneActBench: Can Agents Act on the 3D Scenes They See?
arXiv cs.AI · Опубликовано · 3 мин чтения
За 30 секунд
- Что произошло
- SceneActBench is a benchmark testing whether vision-language model agents can perform multi-object actions on 3D scenes, not just describe them.
- Почему это важно
- Matters for engineers building embodied AI systems or agents that must manipulate complex 3D environments based on visual input.
- На что обратить внимание
- Current VLM agents score only 38.6 to 50.2 percent across tasks, with none performing consistently well, indicating significant capability gaps remain.
Послушать это резюме
Из статьи
-->
Computer Science > Artificial Intelligence
arXiv:2607.22393v1 (cs)
[Submitted on 24 Jul 2026]
Title: SceneActBench: Can Agents Act on the 3D Scenes They See?
Authors: Yifei Zhao , Xiangxin Zhou , Wenhao Yang , Jiaqi Tang , Pu Jian , Huanjin Yao , Jiarui Yao , Haowei Lin , Chunchao Guo , Zhuo Chen , Wenkai Lyu , Jianzhu Ma , Xueqian Wang , Wenxi Zhu
View a PDF of the paper titled SceneActBench: Can Agents Act on the 3D Scenes They See?, by Yifei Zhao and 13 other authors
View PDF HTML (experimental)
Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38
Фрагмент оригинала. Полный текст читайте в первоисточнике.
- agent
- language model
The Agent Architect
Один паттерн, один компромисс, одна история сбоя в продакшене. Короткий еженедельный брифинг для тех, кто строит агентные системы.
Одно письмо в неделю, отписка в один клик. Адрес используется только для рассылки брифинга.