Loading patterns…
Test-Time Scaling(TTS)
Allocating additional inference-time computation to candidate generation, search, verification, or refinement
In 30 seconds
- What
- Generates multiple candidate solutions at inference time, then ranks or filters them using verifiers before returning the best result.
- When to use
- High-stakes outputs where quality matters more than latency, and you can afford multiple forward passes to find better answers.
- Watch out
- Compute costs multiply by candidate count; verification overhead can exceed generation cost if verifiers are expensive or poorly targeted.
Ask the AI expert about this pattern
Opens the assistant with your question prefilled. You review it before sending.
Test-Time Scaling: Overview
Allocating additional inference-time computation to candidate generation, search, verification, or refinement
- Inference-time compute budgets
- Best-of-N candidate sampling
- Search over reasoning paths
- Verifier-guided selection
- Iterative critique and refinement
- Adaptive stopping rules
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.
References
The papers, specifications, and repositories this pattern is based on.
From the engineer behind this catalog
Get your agent architecture reviewed
This page documents one pattern. Your system runs dozens, and most failures live in how they fit together. Have the whole design reviewed against the 288 patterns in this catalog: architecture, reliability, evaluation and cost, every finding mapped to the pattern that fixes it.
€750 instead of €1,500, one week, written report and walkthrough call, until 30 September