Dans l'actualité
vLLM x Novita AI: Chord, Faster INT4 MoE for Kimi K2.x. Up to 1.3x on H200, 2.15x on Untuned B300
vLLM · Novita AI and the vLLM Team · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- Novita AI open-sourced Chord, a W4A16 MoE CUDA operator for INT4 quantized weights, achieving 1.3x speedup on H200 and 2.15x on B300.
- Pourquoi ça compte
- Matters for engineers serving Kimi K2.x models or other INT4 MoE systems on Hopper and Blackwell GPUs seeking inference performance gains.
- Vigilance
- Grouped kernel integration with vLLM remains work-in-progress; indexed path is production-ready but requires explicit quantization humming selection to avoid automatic backend fallback.
- llm
- kernel
- cuda
- serving
- vllm
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.