
Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo
Moonshot's release of the largest open-weight model ever — under a near-permissive license, with day-zero third-party hosting — resets the open/closed gap and hands every enterprise a frontier-class model it can run itself.
On July 27, Moonshot AI did the thing that cannot be reversed: it published the full weights of Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, to Hugging Face for anyone on earth to download. It is the largest open-weight model ever released, and as Nathan Lambert argued in a widely-read analysis, it is best understood not as a product launch but as an escalation.
The specifications
K3 activates roughly 104 billion of its 2.8 trillion parameters per token, supports a 1-million-token context window, and handles text, images and video natively. Moonshot introduced a new attention mechanism, Kimi Delta Attention, which it claims makes long-context inference up to six times cheaper than prior approaches. Using MXFP4 weight quantization, the full model compresses to roughly 1.4TB, with smaller distributions starting around 594GB.
On benchmarks, K3 took first place across six of seven domains in the Frontend Code Arena, scored 88.3 on Terminal-Bench 2.1, and — per Moonshot's own candid technical report — matches Claude Fable 5 on coding tasks at about a third of the cost, though running roughly four times slower. It is not the best model in the world. It is close enough to the best, and free.
The license is the weapon
The strategic core is the license: a Modified MIT that is permissive enough for commercial use and modification. Combined with day-zero third-party hosting — cloud providers were serving the model within hours of the drop — this collapses the friction that historically limited open weights. An enterprise no longer trades quality for control; it can have frontier-class capability and full data sovereignty.
That is what makes the release an escalation. Each time a Chinese lab ships open weights this good under a license this loose, it lowers the price of frontier AI toward zero and raises the pressure on closed labs whose business depends on the gap. K3 follows Qwen, GLM, DeepSeek and Hunyuan in a sustained campaign that has made "open and Chinese" the default answer for cost-sensitive deployment worldwide.
The catch: you still need a data center
The democratization has a hard ceiling. Self-hosting K3 requires roughly 64 NVIDIA H100 or B200 GPUs across eight servers — accessible only to well-resourced organizations. The weights are free; the iron to run them is not. In practice, most users will consume K3 through hosted APIs, which reintroduces exactly the dependency openness was supposed to remove, just with more vendors. The true beneficiaries of the "free" model are the clouds that rent the hardware to run it.
The unresolved shadow
K3 ships under the cloud of the White House's accusation that Moonshot built it by covertly distilling Anthropic's Fable — a claim no benchmark can settle and the released weights cannot dispel. That tension now moves from theoretical to permanent: the model is downloaded, forked and deployed globally, and forensic researchers armed with full weight access will spend months producing verdicts both Washington and Beijing will quote selectively.
For the open-weights movement, though, the provenance fight is almost beside the point. The weights are out. They cannot be recalled. Whatever the courts and policy shops eventually decide, the capability is now permanent global infrastructure — and that irreversibility, more than any benchmark score, is the story of Kimi K3.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake

SciCodePile: A 128GB Corpus and a Brutal Benchmark Where Top Models Solve 12% of Scientific Code Tasks
