In the news
OSWorld-Pro: Process-based Evaluation for Computer Use Agents
arXiv cs.AI · Published · 3 min read
In 30 seconds
- What happened
- OSWorld-Pro introduces over 300 tasks with 2800 subgoals and 67000 human annotations to evaluate computer-use agents step-by-step rather than only final outcomes.
- Why it matters
- Engineers building or testing AI agents that interact with graphical interfaces need granular failure analysis beyond pass-fail metrics to guide improvements.
- Watch out
- Top models like Claude Opus 5 score 75.7% on OSWorld-Pro versus 83.4% on OSWorld, suggesting the benchmark is substantially harder and may not reflect real-world difficulty.
- agent
- eval
- phi
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.