In the news
SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 Lightning
LMSYS · Published · 3 min read
In 30 seconds
- What happened
- SGLang now supports NVIDIA Nemotron 3.5 Lightning, a 30-billion-parameter hybrid mixture-of-experts model with only 3 billion active parameters per token.
- Why it matters
- Engineers building agent systems, local assistants, or specialized enterprise workflows who need efficient inference with reasoning control and multi-token prediction.
- Watch out
- DFlash requires a compatible draft checkpoint and pipeline parallel size 1. DSpark performs best on DGX Spark. Maximum throughput currently achieved without speculative decoding.
Listen to this summary
- nemotron
The patterns behind this
- Energy-Efficient Inference
- Hybrid Secret & Cache Management Pattern
- Local-Distant Agent Data Protection Pattern
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.