ニュース
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
arXiv cs.AI · 公開日 · 読了3分
30秒で要点
- 何が起きたか
- Researchers tested whether agent optimization methods maintain gains when applied repeatedly to new tasks over time using Terminal-Bench 2.0.
- なぜ重要か
- Matters for engineers deploying agents in production where continuous optimization happens as new failures and tasks emerge.
- 注意点
- Only RELAI-VCL maintained and improved gains across optimization rounds; other methods either degraded or plateaued, suggesting most gains may not be stable.
この要約を音声で聴く
記事より
-->
Computer Science > Artificial Intelligence
arXiv:2607.14004v1 (cs)
[Submitted on 15 Jul 2026]
Title: Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Authors: Wenxiao Wang , Priyatham Kattakinda , Soheil Feizi
View a PDF of the paper titled Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0, by Wenxiao Wang and 2 other authors
View PDF HTML (experimental)
Abstract: Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again on newly arrived tasks without eroding the gains the first round produced? We study this question with a two-phase continual-learning evaluation built from hard tasks in Terminal-Bench 2.0, comparing three approaches to agent-harness optimization (GEPA, Meta Harness, and RELAI's Verifiable Continual Learning, RELAI-VCL) under identical optimization budgets. All three methods improve over the baseline agent in the conventional, static, single-phase
原文からの抜粋です。全文は配信元でお読みください。
- agent
The Agent Architect
1つのパターン、1つのトレードオフ、1つの本番障害事例。エージェントシステムを構築する人のための短い週刊ブリーフィング。
週1回のメール、ワンクリックで購読解除できます。アドレスはブリーフィングの送信のみに使用します。