SGLang added immediate support for Muse Glimmer, a multimodal model designed for local agentic workflows.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →Tencent Hunyuan contributes high-performance attention, router GEMM, and MoE kernels to SGLang.
SpecForge v0.3.0 releases a unified disaggregated and colocated speculative decoding stack with new open SpecBundle draft models.
End-to-end reinforcement learning with 8-bit and 4-bit formats on Blackwell hardware.
SGLang improves quantization stack implementation.
LMSYS and Miles add day-0 Kimi K3 support.
SGLang and Miles added day-0 support for Inkling, a frontier multimodal model.
SGLang achieved 500 tokens per second serving GLM5.2 NVFP4 agentic workloads in two weeks.
DeepSeek-V4 Flash RL training now runs on AMD Instinct MI355X GPUs with Miles.
DSpark implements speculative decoding with confidence-driven variable-length verification in SGLang.
Agent-assisted development tools explored for SGLang.
MOSS-TTS Local Transformer v1.5 serves native-streaming 48 kHz speech on SGLang-Omni.
Ling-2.6-1T optimized on TPU with SGLang-JAX hides MoE data movement behind compute.
DFlash and Spec V2 represent next generation speculative decoding techniques.
LMSYS announced the recipient of the 2026 PhD Fellowship.
LMSYS published analysis of token-in-token-out behavior in Miles system.
Higgs Audio v3 TTS integrated with SGLang-Omni for real-time speech generation.
SGLang and Miles added day-zero support for NVIDIA Nemotron 3 Ultra.
Heterogeneous CPU and GPU disaggregation improves VLM serving performance.
AMD Instinct MI355X achieves cost-competitive distributed inference via SGLang with MoRI.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.