AgentBench2023
8 distinct environments, multi-turn.Breadth: whether one agent holds up across unrelated environments.
- Reach for it when
- You are choosing between models and want one number that is not domain-specific.
- What a good score does not prove
- Breadth is not depth. An agent can place well across eight environments and still fail the one workflow you are shipping, because none of the eight is yours.
AgentBench: Evaluating LLMs as Agents (Liu et al., ICLR 2024)arXiv:2308.03688