Chargement des modèles…
OSWorld
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
Aperçu en 30 secondes
- Quoi
- Runs an agent on a real desktop and scores the state of the machine afterwards, across applications, files and the operating system underneath them.
- Quand l'utiliser
- Your agent is expected to operate software through the same interface a person gets, rather than through an API built for it.
- Vigilance
- Failures cluster in knowing where to click rather than what to do, and an end state reached by a route no operator would sanction still scores as success.
Interrogez l'expert IA sur ce pattern
Ouvre l'assistant avec votre question préremplie. Vous la relisez avant l'envoi.
OSWorld: Vue d’ensemble
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
- Real operating systems, not simulations of them
- Execution-based scoring against the final machine state
- Tasks spanning several applications and the filesystem
- Open-ended workflows with no single correct action sequence
- Long-horizon variant (OSWorld 2.0) alongside the original short tasks
- Human baseline collected on the same environments
Recevez le guide de terrain Agent Evals
Les 25 méthodes d’évaluation d’agents condensées en un guide : quel benchmark mesure quoi, quand un score public vous induit en erreur, et comment construire vos évaluations à partir de vos propres échecs. Le lien arrive avec votre confirmation, avec l’hebdo The Agent Architect.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.
Aussi appelé : Computer use benchmark, Desktop agent benchmark, GUI agent evaluation
Références
Les articles, spécifications et dépôts sur lesquels repose ce modèle.
Par l’ingénieur derrière ce catalogue
Voyez ce que vos évaluations laissent passer
Mesurer un agent est plus dur que le livrer, et la plupart des suites restent au vert pendant que la production dérive. Faites relire tout votre dispositif d’évaluation : ce que vous mesurez aujourd’hui, ce que vous ne voyez pas encore, et les régressions que votre suite actuelle laisserait passer.
750 € au lieu de 1 500 €, une semaine, rapport écrit et appel de restitution, jusqu’au 30 septembre