In the news
HPC-Ops × SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent Hunyuan
LMSYS · Published · 3 min read
In 30 seconds
- What happened
- Tencent Hunyuan released HPC-Ops, optimized kernels for attention, routing, and MoE inference now integrated into SGLang's main branch.
- Why it matters
- Engineers deploying large language models on NVIDIA Hopper GPUs who need faster MoE model serving with mixed-length sequences and FP8 quantization.
- Watch out
- Kernels target SM90 Hopper GPUs specifically; performance gains vary by batch size, sequence length distribution, and quantization precision used.
Listen to this summary
- attention
- kernel
The patterns behind this
- Machine Learning Model-Based Routing
- Mixed-Initiative Interface Patterns
- Context Editing & Tool-Result Clearing
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.