NVFP4 KV cache accelerates long-context and agentic inference.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followSGLang and Miles add day-zero support for DeepSeek-V4.1.
DeepSeek-V4-Flash and Kimi-K3 models can run on consumer hardware using SSD Expert Pack.
MiniMax-H3 achieves 1.95x lossless speedup on 8 H200 GPUs, scaling to 6.24x at reduced precision.
SGLang provides day-zero support for Qwen3.8-Flash-Next model.
Ling-3.0-flash uses speculative decoding on Blackwell to optimize batch-1 inference performance.
Mooncake converts fragmented rollout data into efficient bulk I/O operations.
LMSYS examined serving performance limits for DeepSeek-V4-Pro models.
LMSYS released Miles v0.1 for production-level post-training of language models.
Advanced CUDA Graph Techniques in Inference discusses optimization methods for model inference.
SGLang and Miles add day-0 support for Qwen3.8 model.
SGLang added immediate support for NVIDIA Nemotron 3.5 Lightning.
Tencent Hunyuan contributes high-performance attention, router GEMM, and MoE kernels to SGLang.
SpecForge v0.3.0 releases a unified disaggregated and colocated speculative decoding stack with new open SpecBundle draft models.
End-to-end reinforcement learning with 8-bit and 4-bit formats on Blackwell hardware.
SGLang improves quantization stack implementation.
SGLang achieved 500 tokens per second serving GLM5.2 NVFP4 agentic workloads in two weeks.
DSpark implements speculative decoding with confidence-driven variable-length verification in SGLang.
Agent-assisted development tools explored for SGLang.
Ling-2.6-1T optimized on TPU with SGLang-JAX hides MoE data movement behind compute.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.