NVIDIA introduced Nemotron 3.5 ASR Streaming for automatic speech recognition.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →Laguna S 2.1 enables repository-scale game transformation.
Baseten demonstrates fine-tuning Qwen3-TTS for high-quality voice cloning.
Baseten explains model progression from GPT-2 to Kimi K3.
Baseten describes routing Kimi K3 through different harnesses using Baseten Switch.
Announcing Baseten for Model Labs.
Baseten achieves 18x faster tokenization for Kimi K3 million-token workloads.
Baseten provides day-0 API implementation guide for Kimi K3.
Baseten built a faster API for GLM-5.2.
Baseten introduced GLM-5.2 Fast, a faster version of the GLM-5.2 model.
Article discusses methods to optimize LLM inference speed and reduce production costs.
Baseten enables real-time video generation inference.
NVIDIA Nemotron 3 Embed provides fast and accurate retrieval capabilities.
Step 3.7 Flash enables multimodal reasoning at scale.
NVIDIA Nemotron 3 Ultra integrates with LangChain Deep Agents on Baseten.
GLM-5.2 can run in any inference harness on Baseten.
Article explains the difference between AI training and inference.
Live draft model training enables speculative decoding on Baseten.
NVIDIA BioNeMo Agent Toolkit is available on Baseten platform.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.