Dans l'actualité
Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
Together AI · Publié le · 3 min de lecture
En 30 secondes
- Ce qui s'est passé
- Kimi K3 open-weight model matches Claude Fable 5 on DeepSWE coding benchmark at one-third the cost per task.
- Pourquoi ça compte
- Engineering teams evaluating models for software engineering tasks should consider this when cost efficiency and retry tolerance matter more than single-attempt reliability.
- Vigilance
- Kimi K3 is less reliable on individual attempts but excels with multiple retries; Claude Fable 5 remains steadier on first try and solves more tasks consistently.
Écouter ce résumé
Extrait de l'article
Key Takeaways
Kimi K3 matches Claude Fable 5 on quality, costs a third as much per solved task, and as an open model gives teams full control over their deployment.
Kimi K3 vs Claude Fable 5 is close on DeepSWE pass@1: Fable leads 69.9% to 68.5%, a 1.4 point gap.
Give the models more attempts and Kimi K3 pulls ahead. It wins pass@2 (82.0 vs 80.2) and pass@4 (89.4% vs 88.5%).
Kimi K3 is far cheaper: \$4.65 per rollout vs \$13.41, and 2.8x more solved tasks per dollar.
Claude Fable 5 is the more reliable model: it solves more tasks four-for-four (58 vs 45).
Available now · US-hosted
Run Kimi K3 on Together AI
Full 1M context, automatic prefix caching, OpenAI-compatible API, served from US infrastructure.
Open the playground
In our Kimi K3 vs Claude Fable 5 comparison on DeepSWE, a benchmark that tests a model's software engineering capabilities across many task types and programming languages, the most interesting model this month is not the one at the top of the leaderboard. It is Kimi K3, the new open-weight model parked 1.4 points behind Claude Fable 5, at a third of the price.
Kimi K3 landed in DeepSWE on July 16, 2026, with 452 graded rollouts at max effort: 113 real, long-horizon feature requests from live open-source repos, four trials each, graded pass/fail by a hidden test suite. We analyzed all of them against Claude Fable 5 at its best setting (xhigh), An
Extrait de l'original. Lisez l'article complet à la source.
- claude
- kimi
The Agent Architect
Un pattern, un compromis, une panne de production racontée. Un brief hebdomadaire court pour ceux qui construisent des systèmes agentiques.
Un email par semaine, désinscription en un clic. Votre adresse ne sert qu'à envoyer le brief.