The Agent Architect · 2026-W41
The Agent Architect #41: WebArena Evaluation Suite
プレビュー:この号は火曜08:00 UTCに配信されます。
最新号を音声で聴く · 7 min
今週のパターン
WebArena Evaluation Suite
- 概要:
- 実在するウェブサイトのサンドボックス複製上でウェブエージェントを動かし、コードリポジトリ、EC、ソーシャルメディア、業務ワークフローにわたるタスク成功率を測定する。
- 使いどころ:
- 本番導入の前に、ウェブエージェントが現実的な多段階のやり取りをこなせるかどうかを、再現可能で定量的な根拠として示す必要があるとき。
- 注意点:
- セットアップの負担が大きく、評価の実行にも時間がかかる。動的コンテンツ、ボット対策、エージェントが学習していないレイアウトを持つ実サイトには、結果がそのまま当てはまらないこともある。
今週のエージェントAI
- A model guide for the GPT-6 familyOpenAI
OpenAI published guidance for startups on selecting GPT-6 models, adjusting reasoning effort, improving prompts, coordinating tools, and preparing production workflows.
- 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic ScenesarXiv cs.AI
4DCodeBench benchmark evaluates agents on reconstructing dynamic scenes from video as executable graphics programs.
- Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMsarXiv cs.AI
Audio encoder pruning in speech-language models reduces word error rate unevenly across demographic groups.
- Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agentsHacker News
Magnitude, a Y Combinator S25 startup, launches a self-optimizing inference engine for agents.
- Show HN: OpenAPPA – open-source deterministic guardrails that don't break agentsHacker News
OpenAPPA is an open-source project providing deterministic guardrails for AI agents without breaking their functionality.
The Agent Architect
1つのパターン、1つのトレードオフ、1つの本番障害事例。エージェントシステムを構築する人のための短い週刊ブリーフィング。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。