In the news
SceneActBench: Can Agents Act on the 3D Scenes They See?
arXiv cs.AI · Published · 3 min read
In 30 seconds
- What happened
- SceneActBench is a benchmark testing whether vision-language model agents can perform multi-object actions on 3D scenes, not just describe them.
- Why it matters
- Matters for engineers building embodied AI systems or agents that must manipulate complex 3D environments based on visual input.
- Watch out
- Current VLM agents score only 38.6 to 50.2 percent across tasks, with none performing consistently well, indicating significant capability gaps remain.
Listen to this summary
From the article
-->
Computer Science > Artificial Intelligence
arXiv:2607.22393v1 (cs)
[Submitted on 24 Jul 2026]
Title: SceneActBench: Can Agents Act on the 3D Scenes They See?
Authors: Yifei Zhao , Xiangxin Zhou , Wenhao Yang , Jiaqi Tang , Pu Jian , Huanjin Yao , Jiarui Yao , Haowei Lin , Chunchao Guo , Zhuo Chen , Wenkai Lyu , Jianzhu Ma , Xueqian Wang , Wenxi Zhu
View a PDF of the paper titled SceneActBench: Can Agents Act on the 3D Scenes They See?, by Yifei Zhao and 13 other authors
View PDF HTML (experimental)
Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38
Extract from the original. Read the full piece at the source.
- agent
- language model
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.