Loading...
AgentBench(AgentBench)
The original AgentBench study evaluated its published model roster across 8 diverse environments and multi-turn, open-ended settings.
In 30 seconds
- What
- Runs agents through eight diverse environments (SQL, games, web, OS) with multi-turn interactions to measure real-world capability across domains.
- When to use
- Comparing agent models or validating whether your agent handles varied, open-ended tasks beyond single-domain benchmarks.
- Watch out
- High complexity and computational cost; results may not transfer to your specific use cases or custom environments.
Loading technique guide…