In the news
Day 0 Support for Qwen3.8-2.4T-A95B on vLLM
vLLM · vLLM Team and Inferact · Published · 3 min read
In 30 seconds
- What happened
- vLLM released Day-0 support for Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter sparse MoE model with optimized kernels for NVIDIA and AMD hardware.
- Why it matters
- Engineers deploying large language models need this when running inference on Qwen3.8 across NVIDIA B300 or AMD MI355X clusters with memory constraints.
- Watch out
- The model requires at least two GPU nodes for full precision or one node for FP4 quantized versions; reasoning workloads need high max_tokens allocation for accurate results.
Listen to this summary
- llm
- quantiz
- kernel
- vllm
- qwen
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.