In the news
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
NVIDIA Developer · Published · 3 min read
In 30 seconds
- What happened
- NVIDIA GB300 NVL72 achieved 1,648 TFLOPs per GPU pre-training DeepSeek-V3 671B, a world record for mixture-of-experts model training.
- Why it matters
- Matters for engineers scaling large language models who need to understand communication bottlenecks and co-designed hardware-software performance gains.
- Watch out
- Results use optimized software stacks; real-world performance depends on framework choice, model architecture, and whether your workload matches DeepSeek-V3 characteristics.
- token
- mixture of experts
- deepseek
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.