Omi Iyamu · Personal DossierVol. XVII · 2026 Edition
Omi Iyamu.
← All essays
2026 · 07 · 283 min read

Kimi K3 open weights ship: 2.8T MoE lands a hair behind the closed frontier

# The open-weight frontier is now a download away

Moonshot AI released the full open weights of Kimi K3 on July 27. Not previewed, not partial, not for approved partners only. A 2.8-trillion-parameter mixture-of-experts model, 104 billion active per token, native text, image, and video input, a 1,048,576-token context window, and MXFP4 quantization that lands the download at roughly 1.4 to 1.56 terabytes across 96 shards on Hugging Face. Together AI and Modal shipped day-0 hosting. The license is Moonshot's own, not MIT — read it before you build a business on it.

The numbers matter more than the size. On GDPval-AA v2, the benchmark that tries to measure how well a model does actual knowledge work, Kimi K3 lands at roughly 1668 Elo. That puts it behind Claude Fable 5 at 1760 and GPT-5.6 Sol at 1748, and comfortably ahead of Claude Opus 4.8 at 1600. On SWE-bench Verified it scores 76.8%. On FrontierSWE it beats Sol 81.2 to 71.3. On Program Bench it edges Sol 77.8 to 77.6. Not state of the art on every axis, but within striking distance on most of the ones a working team cares about.

The build-versus-buy math changed on Monday.

Two weeks ago in Brief 64 I wrote about the Hugging Face incident, and about a specific detail that stuck with me: when Hugging Face needed a model it could trust for incident response, it reached for an open-weight model, not a US flagship, because the flagship's guardrails got in the way. Kimi K3 is now the strongest open-weight option most teams have ever had access to. If you have been sitting on a fence — API convenience versus self-hosted control, closed frontier versus open near-frontier — the fence moved.

Here is what I'm doing with it this week.

First, the routing exercise. Every model routing table I run for a client gets a new column: what does this call look like on Kimi K3, at self-hosted cost, at self-hosted latency, with our own eval set? Most calls will still route to a workhorse model like Gemini 3.6 Flash or Claude Haiku 5. But the top-tier calls that were parked on Fable or Sol because Opus 4.8 was not quite good enough — some of those will now route to a self-hosted K3 instance, and the compliance story for that instance will be materially different from the API story.

Second, the data-sovereignty conversation. If you serve EU users, the AI Act's enforcement powers over general-purpose model providers go live August 2, five days from now. Article 50 transparency duties are already strictly enforceable. A self-hosted K3 gives you a story to tell about training-data provenance, incident reporting, and where the weights actually sit that a US-hosted API cannot match. It also gives you a story to tell about a Chinese-trained model, which is a different conversation — one your legal team should be in the room for before you sign the license.

Third, the honest reject. Kimi K3 is not the answer for a small team without infrastructure. 1.4 terabytes of weights, 104 billion active parameters, a serving stack that has to justify itself against a $2 per million token API — the economics only work above a specific call volume. If you are shipping fewer than a few million tokens a day, the API is still cheaper than the pager duty. If you are shipping tens of millions, run the math again this week.

A few caveats worth stating out loud. Kimi's own team reported most of these benchmarks. Independent verification is coming; wait for it before you quote the numbers in a board deck. The license is custom and includes restrictions I would want a lawyer to read before signing. And open weights is not open source — you get the weights, not the training data or the training code, so anyone who cares about reproducibility should keep their expectations calibrated.

The larger pattern is what I keep coming back to. Twelve months ago the open-weight frontier was a full generation behind. Nine months ago it was half a generation behind. Today, on most of the benchmarks a working product team actually runs, it is a hair behind the closed flagships and dropping. If you had a build-versus-buy decision parked in Q3 because the open option was not good enough, take it off the shelf.

I'll write a longer piece next week on what the self-hosting cost curve actually looks like for teams shipping into regulated domains. If you're running the numbers on Kimi K3 for a specific product this week and want a second set of eyes, reply.

If this was useful, the weekly Brief covers shorter ideas like this every Wednesday.
Read the Briefs →
© Omi Iyamu · MMXXVIContact → · linkedin.com/in/omiiyamu