In the news
vLLM Reaches 25K Total TPS/GPU on Qwen3.5
vLLM · vLLM Team · Published · 3 min read
In 30 seconds
- What happened
- vLLM achieved 25,000 tokens per second per GPU serving Qwen3.5 on GB200 NVL72 systems using disaggregated prefill-decode architecture.
- Why it matters
- Matters for engineers deploying large language models at scale who need to maximize inference throughput on multi-GPU clusters with hybrid attention architectures.
- Watch out
- Results use fixed 8K input and 1K output sequence lengths on random data; real-world performance varies with actual workload patterns and sequence length distributions.
Listen to this summary
- llm
- kernel
- serving
- vllm
- qwen
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.