In the news
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
NVIDIA Developer · Published · 3 min read
In 30 seconds
- What happened
- NVIDIA GB300 NVL72 achieved 1,648 TFLOPs per GPU pre-training DeepSeek-V3 671B, a world record for mixture-of-experts model training.
- Why it matters
- Matters for engineers scaling large language models who need to understand communication bottlenecks and co-designed hardware-software performance gains.
- Watch out
- Results use optimized software stacks; real-world performance depends on framework choice, model architecture, and whether your workload matches DeepSeek-V3 characteristics.
Listen to this summary
- token
- mixture of experts
- deepseek
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.