Muster werden geladen…
OSWorld
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
In 30 Sekunden
- Was
- Runs an agent on a real desktop and scores the state of the machine afterwards, across applications, files and the operating system underneath them.
- Wann einsetzen
- Your agent is expected to operate software through the same interface a person gets, rather than through an API built for it.
- Achtung
- Failures cluster in knowing where to click rather than what to do, and an end state reached by a route no operator would sanction still scores as success.
Fragen Sie den KI-Experten zu diesem Pattern
Öffnet den Assistenten mit vorbereiteter Frage. Sie prüfen sie vor dem Senden.
OSWorld: Überblick
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
- Real operating systems, not simulations of them
- Execution-based scoring against the final machine state
- Tasks spanning several applications and the filesystem
- Open-ended workflows with no single correct action sequence
- Long-horizon variant (OSWorld 2.0) alongside the original short tasks
- Human baseline collected on the same environments
Den Agent-Evals-Field-Guide erhalten
Alle 25 Methoden zur Agentenbewertung in einem Leitfaden: welcher Benchmark was misst, wann ein öffentlicher Score in die Irre führt und wie Sie Evals aus Ihren eigenen Fehlern bauen. Der Link kommt mit Ihrer Bestätigung, zusammen mit dem wöchentlichen Agent Architect.
Wöchentliche E-Mail, Abmeldung mit einem Klick. Ihre Adresse wird nur für das Briefing verwendet.
Auch bekannt als: Computer use benchmark, Desktop agent benchmark, GUI agent evaluation
Quellen
Die Veröffentlichungen, Spezifikationen und Repositories, auf denen dieses Muster beruht.
Vom Ingenieur hinter diesem Katalog
Finden Sie heraus, was Ihre Evals übersehen
Einen Agenten zu messen ist schwerer, als ihn auszuliefern, und die meisten Suiten bleiben grün, während die Produktion abdriftet. Lassen Sie Ihr Evaluations-Setup durchgehend prüfen: was Sie heute messen, was Sie noch nicht sehen, und welche Regressionen Ihre Suite derzeit durchlässt.
750 € statt 1.500 €, eine Woche, schriftlicher Bericht und Walkthrough-Call, bis 30. September