GPT-5.6 Luna leads DeepSeek-V4 Flash by 14 points on DeepSWE; DeepSeek delivers 4.8x solves per dollar.
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we follow →GPT-5.6 Luna leads DeepSeek-V4 Flash by 14 points on DeepSWE; DeepSeek delivers 4.8x solves per dollar.
Kimi K3 is an open 3-trillion-parameter model available via Together AI API with benchmarks and code examples provided.
Together AI provides guidance on autoscaling metrics, scale windows, and cold start budgeting for LLM inference endpoints.
Together AI partners with Moonshot AI to serve Kimi models natively.
Together AI Dedicated Model Inference uses endpoints, deployments, and configs as resource model components with capacity-aware routing.
ThunderAgent scheduler achieves over 2x single-node throughput and near-linear multi-node scaling for agentic inference by eliminating KV cache thrashing.
Kimi K3 achieves 2.8x cost efficiency on DeepSWE pass@4 versus GPT-5.6 Sol; routing reaches 85.6% performance.
Kimi K3 achieves 2.8x more code problem solves per dollar than Claude Fable 5 on DeepSWE benchmarks.
Together AI offers production platform for deploying open-weight AI models with performance and cost control.
99.9% uptime requires surviving specific failure domains; providers should clarify what each tier actually guarantees.
Together AI offers Provisioned Throughput for reserved inference capacity on open models with 99% uptime SLA.
Together AI presents nine ICML 2026 papers across full stack at Seoul booth B714.
ParallelKernelBench shows leading LLMs solve under one-third of multi-GPU CUDA kernel tasks.
Kimi K2.7 Code generated landing pages 94% cheaper than Claude Fable 5 with comparable quality.
Together AI obtained ISO 27001:2022 certification for enterprise AI security.
Together AI served MiniMax-M3 with sparse attention and optimized decoding to support one million token context windows.
Together AI built fastest speech-to-text by optimizing full system path beyond GPU inference alone.
Together AI benchmarks coding agents achieving 31% higher throughput than TensorRT-LLM and 76% lower cost than Claude Opus.
Together AI and Pearl Research Labs launch discounted Gemma-4-31B-it-pearl inference using Proof of Useful Work.
Violin is open-source video translation tool combining speech recognition, LLM translation, and text-to-speech.
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.