ニュース
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
arXiv cs.AI · 公開日 · 読了3分
30秒で要点
- 何が起きたか
- Researchers evaluated six Vision-Language Models for detecting geometry clipping bugs in video games using autonomous agents and zero-shot prompting.
- なぜ重要か
- Game QA engineers considering automated visual bug detection should understand VLM capabilities and limitations for this specific anomaly detection task.
- 注意点
- All tested VLMs produced substantial false positives on ambiguous frames like near-contact geometry and partial occlusions, limiting standalone use.
この要約を音声で聴く
記事より
-->
Computer Science > Computer Vision and Pattern Recognition
arXiv:2607.25921v1 (cs)
[Submitted on 28 Jul 2026]
Title: Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
Authors: Carlos Celemin , Benedict Wilkins , Adrián Barahona-Ríos , Saman Zadtootaghaj , Nabajeet Barman
View a PDF of the paper titled Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA, by Carlos Celemin and 4 other authors
View PDF HTML (experimental)
Abstract: In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without manual annotation. We benchmark six recent VLMs (Gemini, GPT, Qwen, Gemma, Llama, and Ministral) under a zero-shot prompting setting and analyse their sensitivity to four prompt variants.
Our results show that while the VLMs can capture visual cues associated with geometry clipping, they all produce substantial false positives on visually ambiguous frames such as near-contact geometry and partial occlusions. Gemini-3.1-Flash achi
原文からの抜粋です。全文は配信元でお読みください。
- agent
- language model
- prompt
- gpt
- gemini
The Agent Architect
1つのパターン、1つのトレードオフ、1つの本番障害事例。エージェントシステムを構築する人のための短い週刊ブリーフィング。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。