Decode Context Parallelism in vLLM achieves 3x higher throughput on long-context agentic workloads.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →vLLM achieved 25K tokens per second per GPU on Qwen3.5-397B using GB200 NVL72 disaggregated serving.
vLLM enables inference performance optimizations on Arm CPUs.
vLLM adds P-EAGLE, DFlash, and DSpark parallel drafting algorithms for faster speculative decoding.
vLLM delivers Kimi K3 serving with hybrid prefix caching and DSpark speculative decoding.
vLLM released AFD Plugin enabling attention-FFN disaggregation for mixture-of-experts model serving.
vLLM optimized GLM-5.2 serving on 24 B300 GPUs, reducing mean time per output token from 40ms to 17ms.
vLLM added production-scale Kimi K3 support with optimizations for caching and MoE.
vLLM Semantic Router expanded to build, evaluate, and run Mixture-of-Models systems.
vLLM maintains production quality through extensive CI, nightly benchmarking, and a two-week release cycle.
vLLM supports TML Inkling, a 1T-parameter multimodal model, achieving 380 tokens per second on GB200 GPUs.
vLLM V1 integrates TileRT as a specialized decode engine for latency-critical serving without modifying vLLM.
AMD Quark trains and serves EAGLE3 speculative decoding with vLLM on AMD Instinct GPUs, achieving up to 2.00x throughput gains.
vime now supports ROCm for end-to-end RL post-training on AMD Instinct MI355X GPUs.
Tencent Hunyuan integrated HPC-Ops attention and MoE backends into vLLM for improved latency on NVIDIA H20.
vLLM-Omni serves Qwen3-Omni using staged Thinker-Talker-Code2Wav execution with batching, CUDA Graphs, and async optimization.
vLLM Semantic Router enables micro-agents to match frontier model performance through model API collaboration.
vLLM-Omni optimizes text-to-speech inference for multiple models using staged serving, batching, and CUDA Graphs.
vLLM Semantic Router Fusion runs multiple models with a judge to synthesize answers.
vLLM serves MiniMax M3 with sparse attention, multimodal parsing, and MXFP8 weights for long-context deployment.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.


