In the news
How SWE-Serve Exposes the Gap Between Local Tests and Live Serving
NVIDIA Developer · Published · 3 min read
In 30 seconds
- What happened
- SWE-Serve benchmark reveals that AI coding patches pass local tests 69.4% of the time but only 45.9% when live serving checks are included.
- Why it matters
- Engineers building inference-serving systems need this when evaluating AI agents for production-grade changes to LLM serving software like SGLang.
- Watch out
- SWE-Serve tests only SGLang on single H100 or CPU; results may not generalize to other inference engines, multi-GPU setups, or different hardware configurations.
- agent
- inference
- serving
- eval
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.