In the news
TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
NVIDIA Developer · Published · 3 min read
In 30 seconds
- What happened
- TensorRT Edge-LLM completed MLPerf Edge Agentic benchmark 6.4x faster than reference, running Qwen3.6-27B on Jetson AGX Thor at 52.33 tokens per second.
- Why it matters
- Matters for engineers deploying multi-turn AI agents on edge devices like robots and vehicles where long-context inference and tool calling efficiency are critical.
- Watch out
- Result uses specialized techniques including NVFP4 quantization, tree-based multi-token prediction, and KV cache reuse that require specific hardware and careful configuration to replicate.
- agent
- agentic
- llm
- reasoning
- prompt
The patterns behind this
- Energy-Efficient Inference
- Context Editing & Tool-Result Clearing
- Agentic Context Engineering (Evolving Playbook)
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.