正在加载模式…
OSWorld
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
30秒速览
- 是什么
- Runs an agent on a real desktop and scores the state of the machine afterwards, across applications, files and the operating system underneath them.
- 何时使用
- Your agent is expected to operate software through the same interface a person gets, rather than through an API built for it.
- 注意
- Failures cluster in knowing where to click rather than what to do, and an end state reached by a route no operator would sanction still scores as success.
向AI专家咨询此模式
打开助手并预填您的问题,发送前可先确认。
OSWorld: 概览
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
- Real operating systems, not simulations of them
- Execution-based scoring against the final machine state
- Tasks spanning several applications and the filesystem
- Open-ended workflows with no single correct action sequence
- Long-horizon variant (OSWorld 2.0) alongside the original short tasks
- Human baseline collected on the same environments
The Agent Architect
每周一个模式、一个权衡、一个生产事故案例。为构建智能体系统的人准备的每周简报。
每周一封邮件,一键退订。您的地址仅用于发送简报。
又称: Computer use benchmark, Desktop agent benchmark, GUI agent evaluation
参考资料
该模式所依据的论文、规范和代码仓库。