In the news
Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
NVIDIA Developer · Published · 3 min read
In 30 seconds
- What happened
- NVIDIA Transformer Engine with JAX achieves 10.4x throughput improvement for Mixture of Experts training, reaching 1,068 TFLOPS/GPU on DeepSeek-V3.
- Why it matters
- Matters for engineers training large MoE models at scale who need to maximize GPU utilization and reduce inter-GPU communication overhead.
- Watch out
- Dropless MoE requires specialized kernels for variable token counts; standard capacity-based MoE may remain simpler for some use cases despite efficiency tradeoffs.
- mixture of experts
- qwen
- deepseek
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.