In the news
Introducing preemptible compute: the same compute, half the price
Together AI · Published · 3 min read
In 30 seconds
- What happened
- Together AI released preemptible GPU compute at 50% of on-demand rates for Kubernetes clusters, with five-minute graceful shutdown windows.
- Why it matters
- Engineers running interruptible workloads like experiments, batch jobs, and inference bursts who can checkpoint and retry work.
- Watch out
- Preemptible nodes draw from spare capacity only, so allocated capacity may stay below your requested target during high demand.
The patterns behind this
Each one covers how the technique works, when it earns its cost, and where it breaks.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.