In den Nachrichten
Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
Together AI · Veröffentlicht am · 3 Min. Lesezeit
In 30 Sekunden
- Was passiert ist
- Kimi K3 open-weight model matches Claude Fable 5 on DeepSWE coding benchmark at one-third the cost per task.
- Warum es zählt
- Engineering teams evaluating models for software engineering tasks should consider this when cost efficiency and retry tolerance matter more than single-attempt reliability.
- Achtung
- Kimi K3 is less reliable on individual attempts but excels with multiple retries; Claude Fable 5 remains steadier on first try and solves more tasks consistently.
Diese Zusammenfassung anhören
Aus dem Artikel
Key Takeaways
Kimi K3 matches Claude Fable 5 on quality, costs a third as much per solved task, and as an open model gives teams full control over their deployment.
Kimi K3 vs Claude Fable 5 is close on DeepSWE pass@1: Fable leads 69.9% to 68.5%, a 1.4 point gap.
Give the models more attempts and Kimi K3 pulls ahead. It wins pass@2 (82.0 vs 80.2) and pass@4 (89.4% vs 88.5%).
Kimi K3 is far cheaper: \$4.65 per rollout vs \$13.41, and 2.8x more solved tasks per dollar.
Claude Fable 5 is the more reliable model: it solves more tasks four-for-four (58 vs 45).
Available now · US-hosted
Run Kimi K3 on Together AI
Full 1M context, automatic prefix caching, OpenAI-compatible API, served from US infrastructure.
Open the playground
In our Kimi K3 vs Claude Fable 5 comparison on DeepSWE, a benchmark that tests a model's software engineering capabilities across many task types and programming languages, the most interesting model this month is not the one at the top of the leaderboard. It is Kimi K3, the new open-weight model parked 1.4 points behind Claude Fable 5, at a third of the price.
Kimi K3 landed in DeepSWE on July 16, 2026, with 452 graded rollouts at max effort: 113 real, long-horizon feature requests from live open-source repos, four trials each, graded pass/fail by a hidden test suite. We analyzed all of them against Claude Fable 5 at its best setting (xhigh), An
Auszug aus dem Original. Den vollständigen Text bei der Quelle lesen.
Den vollständigen Artikel lesen
- claude
- kimi
The Agent Architect
Ein Pattern, ein Tradeoff, eine Produktionspanne. Ein kurzes wöchentliches Briefing für alle, die agentische Systeme bauen.
Wöchentliche E-Mail, Abmeldung mit einem Klick. Ihre Adresse wird nur für das Briefing verwendet.