Chargement des modèles…
OSWorld
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
Aperçu en 30 secondes
- Quoi
- Runs an agent on a real desktop and scores the state of the machine afterwards, across applications, files and the operating system underneath them.
- Quand l'utiliser
- Your agent is expected to operate software through the same interface a person gets, rather than through an API built for it.
- Vigilance
- Failures cluster in knowing where to click rather than what to do, and an end state reached by a route no operator would sanction still scores as success.
Interrogez l'expert IA sur ce pattern
Ouvre l'assistant avec votre question préremplie. Vous la relisez avant l'envoi.
OSWorld: Vue d’ensemble
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
- Real operating systems, not simulations of them
- Execution-based scoring against the final machine state
- Tasks spanning several applications and the filesystem
- Open-ended workflows with no single correct action sequence
- Long-horizon variant (OSWorld 2.0) alongside the original short tasks
- Human baseline collected on the same environments
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.
Aussi appelé : Computer use benchmark, Desktop agent benchmark, GUI agent evaluation
Références
Les articles, spécifications et dépôts sur lesquels repose ce modèle.