LLM Model Comparison
Context window, maximum output and token pricing for the models currently worth considering. Every row links to the provider page the numbers came from, because this is the kind of table that goes stale quietly.
Figures verified . Providers change prices and limits without notice; check the source link before committing to a number.
Anthropic
| Model | Tier | Context | Max output | Price / 1M in, out | Source |
|---|---|---|---|---|---|
| Claude Fable 5claude-fable-5 | Frontier | 1M | 128K | $10.00 / $50.00 per 1M tokens | docs |
| Claude Opus 4.8claude-opus-4-8 | Frontier | 1M | 128K | $5.00 / $25.00 per 1M tokens | docs |
| Claude Sonnet 5claude-sonnet-5 | Balanced | 1M | 128K | $2.00 / $10.00 per 1M tokens | docs |
| Claude Haiku 4.5claude-haiku-4-5 | Fast | 200K | 64K | $1.00 / $5.00 per 1M tokens | docs |
OpenAI
DeepSeek
| Model | Tier | Context | Max output | Price / 1M in, out | Source |
|---|---|---|---|---|---|
| DeepSeek V4 Propreviewdeepseek-v4-pro | Open weights | 1M | not stated | self-hosted | docs |
Alibaba
| Model | Tier | Context | Max output | Price / 1M in, out | Source |
|---|---|---|---|---|---|
| Qwen3.6 35B-A3BQwen/Qwen3.6-35B-A3B | Open weights | 262K | not stated | self-hosted | docs |
Meta
| Model | Tier | Context | Max output | Price / 1M in, out | Source |
|---|---|---|---|---|---|
| Llama 4 Maverickllama-4-maverick | Open weights | 1M | not stated | self-hosted | docs |
Reading this table
Context window is a ceiling, not a budget. Attention cost grows with sequence length, so a million-token window is an option you pay for by the token rather than a free allowance. Most production agents are cost-bound long before they are context-bound.
Output price is the one that bites. It is typically several times the input price, and agent loops resend their whole transcript each step. A verbose tool result is charged as input on every subsequent step of the run.
Tier beats vendor. The useful comparison is frontier against frontier and fast against fast. Comparing a frontier model to another vendor's fast one measures your choice of row, not the models.
For how these models differ inside rather than on a spec sheet, see model architectures, and for choosing between approaches rather than models, the Eval Lab runs them side by side with real latency and cost.