In den Nachrichten
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
arXiv cs.AI · Veröffentlicht am · 3 Min. Lesezeit
In 30 Sekunden
- Was passiert ist
- Desktop-Delta Bench introduces a benchmark with 2,013 instances to test whether computer-use AI models can understand GUI state changes from desktop actions.
- Warum es zählt
- Matters for engineers building desktop automation agents who need to verify models correctly interpret action consequences and recover from failures.
- Achtung
- Best models achieve only 65% accuracy on temporal ordering tasks; the benchmark reveals systematic weaknesses but doesn't yet show how to fix them.
Diese Zusammenfassung anhören
Aus dem Artikel
-->
Computer Science > Artificial Intelligence
arXiv:2607.26041v1 (cs)
[Submitted on 28 Jul 2026]
Title: Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Authors: Abhishek Pillai , Samir Kumar Nayak , Yuan Chen
View a PDF of the paper titled Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?, by Abhishek Pillai and 2 other authors
View PDF
Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50 task domains. DDB trajectories targets 3 failure dimensions- state verification, source tracking, and context-aware control- through 2 complementary tasks: 463 3-frame temporal-ord
Auszug aus dem Original. Den vollständigen Text bei der Quelle lesen.
Den vollständigen Artikel lesen
- agent
- inference
The Agent Architect
Ein Pattern, ein Tradeoff, eine Produktionspanne. Ein kurzes wöchentliches Briefing für alle, die agentische Systeme bauen.
Wöchentliche E-Mail, Abmeldung mit einem Klick. Ihre Adresse wird nur für das Briefing verwendet.