vLLM implements Gumbel-max text watermarking using GPU kernels with statistical detection and speculative decoding support.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followvllm-metal brings vLLM's serving stack to Apple Silicon with lower time-to-first-token under concurrent load.
vLLM achieves 5K throughput and 180ms interactivity serving Qwen3.8-2.4T on GB300 NVL72 with PD serving.
vLLM demonstrates multi-GPU video captioning scaling using NVIDIA hardware video decoders.
Speculators and Mooncake enabled multi-node distributed training of Kimi K3 on GB300 NVL72 GPUs.
Novita AI open-sourced Chord, a W4A16 MoE CUDA kernel achieving 1.3x speedup on H200 and 2.15x on B300 for Kimi K2.x models.
vime and RL-Kernel achieve bit-exact consistency between Megatron training and vLLM rollout on AMD MI300X across 200 GRPO steps.
Kimi K3 serving optimizations achieved 2.8x throughput improvement through scheduling, caching, state recovery, and kernel enhancements.
Performance optimization method for LLM serving on AMD Instinct MI355X using bottleneck analysis and queue inspection.
vLLM introduced tiered KV cache offloading across host memory, filesystems, and remote storage to increase serving capacity.
vLLM optimizes KV cache and scheduling for agentic workloads with 14.6x-106x cost advantage.
Tenstorrent accelerators integrated into vLLM via plugin with phase-based scheduling, single-process data parallelism, on-device sampling, and async decode overlap.
vLLM integrates HiSparse memory offloading to allow GLM 5.3 decoding to continue when KV cache exceeds GPU memory.
vLLM-Omni optimizes MiniMax H3 and integrates FastVideo's FastH3 for video generation faster than real-time playback.
vLLM supports speculative decoding on AMD GPUs using draft-and-verify methods including MTP, EAGLE-3, DFlash, and DSpark.
vLLM implements native sharded weight transfer using Ray Direct Transport, transferring Kimi K2 model across 48 nodes in 7.53 seconds.
IsoExec unifies numerical execution across vLLM and Megatron runtimes, reducing trainer-inference logprob mismatch below 1e-6.
Distributed Layerwise Offload in vLLM-Omni serves 124 GB DiT models on 64 GB HBM.
vLLM's DSpark verification uses per-request confidence to adjust draft-verification budget, maintaining throughput-latency tradeoffs across batch sizes one to 256.
vLLM adds day-0 support for Qwen3.8-2.4T-A95B hybrid MoE model with quantized weights on NVIDIA and AMD.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.







