In the news
PD Serving of Qwen3.8-2.4T
vLLM · vLLM Team · Published · 3 min read
In 30 seconds
- What happened
- vLLM achieved 5000 tokens per GPU throughput and 180 generated tokens per user latency serving Qwen3.8-2.4T on GB300 clusters.
- Why it matters
- Engineers optimizing large language model serving performance need reproducible tuning methodologies for disaggregated serving configurations.
- Watch out
- Results are specific to Qwen3.8-2.4T on GB300 hardware; CUDA graph memory estimation often overallocates, reducing actual KV cache availability.
- llm
- serving
- throughput
- vllm
- qwen
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.