Dans l'actualité
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
arXiv cs.AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- Desktop-Delta Bench introduces a benchmark with 2,013 instances to test whether computer-use AI models can understand GUI state changes from desktop actions.
- Pourquoi ça compte
- Matters for engineers building desktop automation agents who need to verify models correctly interpret action consequences and recover from failures.
- Vigilance
- Best models achieve only 65% accuracy on temporal ordering tasks; the benchmark reveals systematic weaknesses but doesn't yet show how to fix them.
Écouter ce résumé
Extrait de l'article
-->
Computer Science > Artificial Intelligence
arXiv:2607.26041v1 (cs)
[Submitted on 28 Jul 2026]
Title: Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Authors: Abhishek Pillai , Samir Kumar Nayak , Yuan Chen
View a PDF of the paper titled Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?, by Abhishek Pillai and 2 other authors
View PDF
Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50 task domains. DDB trajectories targets 3 failure dimensions- state verification, source tracking, and context-aware control- through 2 complementary tasks: 463 3-frame temporal-ord
Extrait de l'original. Lisez l'article complet à la source.
- agent
- inference
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.