The letter

One letter a week, in your inbox.

The week's signal, what shipped, what is worth running tonight, and the editor's note. No tracking, no ads, nothing else.

We keep your address, your language and the date you joined, nothing else. Every letter has a one-click unsubscribe link that deletes the record.

Together AI

18 stories

Together AIToolsHow to train your own Jev for $17 We just launched our own Jev-like classifier, together/Tev1-4B-experimental, on top of Qwen3.5 4B on Together’s serverless platform. In this blog post we’ll show you how to fine-tune your own version!

September 23

Together AIToolsCanary rollouts: upgrade models in production without downtime A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure.

September 22

Together AIToolsHow a global fintech scaled coding agent traffic with Dedicated Model Inference Inside a global bank's shift to self-serve dedicated inference: how Together's DMI gave engineering teams direct control over scaling, models, and testing.

September 18

Together AIToolsMigrating from closed to open source models, Together Moving from closed to open source models can take weeks, not years. A five-stage playbook: discover, evaluate, adapt, decide, and production.

September 16

Together AIToolsTogether AI expands fine-tuning service with more models, live metrics, and finer controls Together Fine-Tuning adds the latest open-weight models, live experiment tracking, Expert LoRA, early stopping, tokenized dataset previews, pre-flight validation, and lower training prices on…

September 11

Together AIToolsIntroducing preemptible compute: the same compute, half the price Together GPU Clusters now supports preemptible compute: the same GPU capacity at a flat 50% of the on-demand rate, with a five-minute drain window.

September 10

Together AIToolsTo Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72! We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL.

September 10

Together AIToolsThe Open Source AI Stack A deep dive into the open model AI stack — model, inference, gateways and routers, harness, and tools — and how keeping each layer independent lets you swap in a new open model in minutes instead of…

September 9

Together AIToolsGLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.

August 21

Together AIToolsGLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.

August 21

Together AIToolsDeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

August 17

Together AIToolsA/B test models in production Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.

August 17

Together AIToolsKimi K3: the complete developer guide Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.

August 1

Together AIToolsTogether AI announces strategic partnership with Moonshot AI to natively serve Kimi models Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.

July 29

Together AIToolsThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node…

July 29

Together AIToolsThe production platform for open-weight AI inference Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.

July 23

Together AIToolsTogether AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.

July 20

Together AIToolsWhat does 99.9% uptime mean for inference? Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference…

July 16