パターンを読み込んでいます…
OSWorld
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
30秒でわかる概要
- 概要
- Runs an agent on a real desktop and scores the state of the machine afterwards, across applications, files and the operating system underneath them.
- 使いどころ
- Your agent is expected to operate software through the same interface a person gets, rather than through an API built for it.
- 注意点
- Failures cluster in knowing where to click rather than what to do, and an end state reached by a route no operator would sanction still scores as success.
このパターンについてAIエキスパートに質問
質問が入力済みの状態でアシスタントが開きます。送信前に内容を確認できます。
OSWorld: 概要
Executable desktop environments where an agent is scored on the state it leaves behind after doing real work across applications, files and the operating system.
- Real operating systems, not simulations of them
- Execution-based scoring against the final machine state
- Tasks spanning several applications and the filesystem
- Open-ended workflows with no single correct action sequence
- Long-horizon variant (OSWorld 2.0) alongside the original short tasks
- Human baseline collected on the same environments
エージェント評価フィールドガイドを受け取る
25のエージェント評価手法を1冊に凝縮:どのベンチマークが何を測るか、公開スコアが誤解を招くのはどんなときか、自分の失敗から評価を組み立てる方法。確認メールと一緒にリンクが届き、週刊The Agent Architectも購読できます。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。
別名: Computer use benchmark, Desktop agent benchmark, GUI agent evaluation
参考文献
このパターンの根拠となる論文、仕様、リポジトリです。
このカタログを作ったエンジニアが担当
評価が見落としているものを可視化
エージェントの評価は実装よりも難しく、多くのテストは緑のまま本番だけがずれていきます。評価の仕組みを端から端まで点検します。いま測れているもの、まだ見えていないもの、そして現状のテストが素通りさせる劣化を洗い出します。
€750(通常€1,500)・1週間・文書レポートとウォークスルーコール・9月30日まで