Together AI released Tev1-4B-experimental, a Jailbreak-detection classifier based on Qwen3.5 4B, available for fine-tuning on their serverless platform.
News Hub
What actually shipped in agent engineering, pulled from the labs, arXiv and Hacker News.
See who we followTogether AI enables canary rollouts for model upgrades using staged traffic ramps, metric gates, and automatic rollback.
A global bank used Together AI's Dedicated Model Inference to give engineering teams direct control over scaling and testing coding agents.
Together AI published a five-stage playbook for migrating from closed to open source models.
Together AI adds more models, live metrics, and controls to its fine-tuning service.
ThunderKittens ported to NVIDIA Vera Rubin NVL72 achieves 22 PFLOPS, competitive with cuBLAS.
Together AI offers preemptible GPU compute at 50% of on-demand rates with five-minute drain windows.
Together AI describes the open source AI stack.
GLM-5.3 Flash achieves 17x lower cost than GLM-5.3 with only 5.6 point pass@1 reduction on DeepSWE.
GLM-5.3 achieves better pass@4 at half the cost of GPT-5.6 Sol on DeepSWE.
GLM-5.3 costs 5.4x less than Claude Fable 5 while matching pass@1 on DeepSWE.
DeepSeek V4 Pro 0813 achieved 82.7% pass rate on DeepSWE cascade routing, outperforming Claude Fable 5 on cost-adjusted coding tasks.
GPT-5.6 Luna leads DeepSeek-V4 Flash by 14 points on DeepSWE; DeepSeek delivers 4.8x solves per dollar.
Together AI provides guidance on autoscaling metrics, scale windows, and cold start budgeting for LLM inference endpoints.
ThunderAgent scheduler achieves over 2x single-node throughput and near-linear multi-node scaling for agentic inference by eliminating KV cache thrashing.
Together AI Dedicated Model Inference uses endpoints, deployments, and configs as resource model components with capacity-aware routing.
Together AI partners with Moonshot AI to serve Kimi models natively.
Kimi K3 achieves 2.8x more code problem solves per dollar than Claude Fable 5 on DeepSWE benchmarks.
Together AI offers production platform for deploying open-weight AI models with performance and cost control.
99.9% uptime requires surviving specific failure domains; providers should clarify what each tier actually guarantees.
The Agent Architect
One pattern, one tradeoff, one production failure story. A short weekly briefing for people building agentic systems.
Weekly email, one-click unsubscribe. We only use your address to send the briefing.




.png&w=160)













.png&w=160)