In den Nachrichten
SceneActBench: Can Agents Act on the 3D Scenes They See?
arXiv cs.AI · Veröffentlicht am · 3 Min. Lesezeit
In 30 Sekunden
- Was passiert ist
- SceneActBench is a benchmark testing whether vision-language model agents can perform multi-object actions on 3D scenes, not just describe them.
- Warum es zählt
- Matters for engineers building embodied AI systems or agents that must manipulate complex 3D environments based on visual input.
- Achtung
- Current VLM agents score only 38.6 to 50.2 percent across tasks, with none performing consistently well, indicating significant capability gaps remain.
Diese Zusammenfassung anhören
Aus dem Artikel
-->
Computer Science > Artificial Intelligence
arXiv:2607.22393v1 (cs)
[Submitted on 24 Jul 2026]
Title: SceneActBench: Can Agents Act on the 3D Scenes They See?
Authors: Yifei Zhao , Xiangxin Zhou , Wenhao Yang , Jiaqi Tang , Pu Jian , Huanjin Yao , Jiarui Yao , Haowei Lin , Chunchao Guo , Zhuo Chen , Wenkai Lyu , Jianzhu Ma , Xueqian Wang , Wenxi Zhu
View a PDF of the paper titled SceneActBench: Can Agents Act on the 3D Scenes They See?, by Yifei Zhao and 13 other authors
View PDF HTML (experimental)
Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38
Auszug aus dem Original. Den vollständigen Text bei der Quelle lesen.
Den vollständigen Artikel lesen
- agent
- language model
The Agent Architect
Ein Pattern, ein Tradeoff, eine Produktionspanne. Ein kurzes wöchentliches Briefing für alle, die agentische Systeme bauen.
Wöchentliche E-Mail, Abmeldung mit einem Klick. Ihre Adresse wird nur für das Briefing verwendet.