Dans l'actualité
SceneActBench: Can Agents Act on the 3D Scenes They See?
arXiv cs.AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- SceneActBench is a benchmark testing whether vision-language model agents can perform multi-object actions on 3D scenes, not just describe them.
- Pourquoi ça compte
- Matters for engineers building embodied AI systems or agents that must manipulate complex 3D environments based on visual input.
- Vigilance
- Current VLM agents score only 38.6 to 50.2 percent across tasks, with none performing consistently well, indicating significant capability gaps remain.
Écouter ce résumé
Extrait de l'article
-->
Computer Science > Artificial Intelligence
arXiv:2607.22393v1 (cs)
[Submitted on 24 Jul 2026]
Title: SceneActBench: Can Agents Act on the 3D Scenes They See?
Authors: Yifei Zhao , Xiangxin Zhou , Wenhao Yang , Jiaqi Tang , Pu Jian , Huanjin Yao , Jiarui Yao , Haowei Lin , Chunchao Guo , Zhuo Chen , Wenkai Lyu , Jianzhu Ma , Xueqian Wang , Wenxi Zhu
View a PDF of the paper titled SceneActBench: Can Agents Act on the 3D Scenes They See?, by Yifei Zhao and 13 other authors
View PDF HTML (experimental)
Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38
Extrait de l'original. Lisez l'article complet à la source.
- agent
- language model
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.