In the news
DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
Apple Machine Learning Research · Published · 3 min read
In 30 seconds
- What happened
- Apple researchers released DeepAmbigQA, a 3,600-question benchmark testing whether LLMs can answer complex questions requiring both name disambiguation and multi-hop reasoning.
- Why it matters
- Engineers building search-augmented LLM systems should care when evaluating whether their models return complete answer sets to ambiguous, multi-step questions.
- Watch out
- Even GPT-5 achieves only 0.13 exact match on ambiguous questions, suggesting current LLMs struggle fundamentally with answer completeness rather than this being a simple benchmark limitation.
Listen to this summary
- llm
- language model
- reasoning
- eval
- benchmark
The patterns behind this
- Proactive Clarification & Active Disambiguation
- Eval-Driven Development (Agent CI)
- Agentic Context Engineering (Evolving Playbook)
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.