Chargement des modèles…
Terminal-Bench(TB)
Hard command-line tasks in isolated environments, each with a human-written solution and tests that decide whether the agent actually finished.
Aperçu en 30 secondes
- Quoi
- Poses hard command line tasks in isolated containers, each carrying a human-written solution and tests that decide afterwards whether the work was actually finished.
- Quand l'utiliser
- You are evaluating agents that do infrastructure, build or debugging work in a shell, and you need correctness decided rather than judged.
- Vigilance
- The score names a model and a harness together: the same weights under a weaker scaffold give up exactly where a stronger one reads the error and recovers.
Interrogez l'expert IA sur ce pattern
Ouvre l'assistant avec votre question préremplie. Vous la relisez avant l'envoi.
Terminal-Bench: Vue d’ensemble
Hard command-line tasks in isolated environments, each with a human-written solution and tests that decide whether the agent actually finished.
- One container per task, so runs cannot contaminate each other
- Every task carries a human-written reference solution
- Verified by tests rather than by inspecting the transcript
- Drawn from real infrastructure, data and debugging work
- Deliberately hard: frontier systems remain well short of solving it
- Harness-agnostic, so the agent and the scaffold are measured together
Recevez le guide de terrain Agent Evals
Les 25 méthodes d’évaluation d’agents condensées en un guide : quel benchmark mesure quoi, quand un score public vous induit en erreur, et comment construire vos évaluations à partir de vos propres échecs. Le lien arrive avec votre confirmation, avec l’hebdo The Agent Architect.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.
Aussi appelé : CLI agent benchmark, Command line agent evaluation
Références
Les articles, spécifications et dépôts sur lesquels repose ce modèle.
Par l’ingénieur derrière ce catalogue
Voyez ce que vos évaluations laissent passer
Mesurer un agent est plus dur que le livrer, et la plupart des suites restent au vert pendant que la production dérive. Faites relire tout votre dispositif d’évaluation : ce que vous mesurez aujourd’hui, ce que vous ne voyez pas encore, et les régressions que votre suite actuelle laisserait passer.
750 € au lieu de 1 500 €, une semaine, rapport écrit et appel de restitution, jusqu’au 30 septembre