In the news
LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook
Liquid AI · Published · 3 min read
In 30 seconds
- What happened
- Liquid AI released DSpark draft models for three LFM2.5 models, enabling speculative decoding that speeds up inference up to 3.2x on GPUs and 2.87x on MacBooks without changing output quality.
- Why it matters
- Matters for engineers deploying language models on GPUs or edge devices who need faster token generation for interactive applications like agentic reasoning or function calling.
- Watch out
- Speedup varies significantly by model size and hardware. The 8B model shows only 1.18x speedup on MacBooks due to current MoE implementation limitations in llama.cpp's Metal backend, requiring future optimization work.
Listen to this summary
- inference
- lfm
Who else ran this
- Up to 3.2x Faster Inference with LFM2.5-DSparkHugging Face
The same event, reported by other publishers we follow.
The patterns behind this
- Function Calling
- Speculative & Parallel Tool Execution
- Agentic Context Engineering (Evolving Playbook)
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.