In the news
Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each
NVIDIA Developer · Published · 3 min read
In 30 seconds
- What happened
- NVIDIA explains when to deploy dense versus Mixture-of-Experts models, showing MoE activates fewer parameters per token for higher throughput.
- Why it matters
- Engineers choosing between model architectures for inference need to understand memory, concurrency, fine-tuning, and quantization tradeoffs specific to their deployment.
- Watch out
- MoE's throughput advantage narrows at high concurrency and latency-sensitive scenarios. Router imbalance during fine-tuning and quantization sensitivity require careful handling.
- throughput
- token
- mixture-of-experts
- nemotron
The patterns behind this
- Parametric Memory
- Proactive Clarification & Active Disambiguation
- Agentic Context Engineering (Evolving Playbook)
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.