Moonshot AI published the weights of Kimi K3, a 2.8 trillion parameter sparse mixture of experts model, on the night of 26 July 2026 US time, hours ahead of its stated 27 July target, making it the largest open weight model ever released.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 4 Tom's Hardware 2026-07-27 Frames Kimi K3 as the largest open weight AI model ever delivered, beating Claude Fable 5 on the Frontend Code Arena benchmark, as China works around US compute limits. Open source The raw four bit weights alone are estimated at roughly 1.4 TB, and running the full model requires a minimum of 8 H100 class GPUs, with Moonshot recommending 64 or more accelerators for serious deployment.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source 2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source The release matters because it moves the open weight frontier from the 1.6 trillion parameter class, where DeepSeek V4 Pro and Meituan's LongCat-2.0 sit, to nearly double that, under a permissive modified MIT license.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source 5 MarkTechPost 2026-07-05 Meituan released LongCat-2.0 on 30 June 2026 under MIT: 1.6 trillion total parameters, roughly 48 billion active per token, native 1 million token context, trained on a 50,000 card domestic Chinese accelerator cluster over 35 trillion plus tokens, SWE-bench Pro 59.5. Open source We assess with high confidence that the release is aimed less at hobbyist self hosters than at the inference platforms, enterprises, and rival labs whose adoption sets defaults for the developer economy, and with moderate confidence that Chinese open weight releases are now the primary force setting the floor under US model pricing.

What actually shipped

K3 is a sparse mixture of experts design: 2.8 trillion total parameters organized into 896 experts, of which 16 activate per token, so the compute paid per token is a small fraction of the headline figure.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source The context window is 1,048,576 tokens, and the model was trained quantization aware from supervised fine tuning onward, shipping in MXFP4 weights with MXFP8 activations, which is why a 2.8 trillion parameter model compresses to roughly 1.4 TB of raw weights rather than several times that.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source Moonshot had run the model as an API only product since its 17 July announcement, at 3 dollars per million input tokens and 15 dollars per million output tokens, before opening the weights ten days later.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source 2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source

The benchmark picture, mostly from Moonshot's own materials and third party leaderboards rather than independent replication, is strong in a specific lane. Amplifi Labs reports 77.8 on Program Bench against 77.6 for GPT-5.6 Sol, 42.0 on SWE Marathon against 40.0 for Opus 4.8, first place on Frontend Code Arena at 1,679 Elo, and 93.5 percent on GPQA.3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source The same guide is candid that K3 trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis.3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source The fair read is a model built, in Amplifi's phrasing, around "tool-calling and terminal use as first-class capabilities" that wins on long horizon agent tasks and frontend generation while remaining a step behind the US frontier on breadth.3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source Those margins over US models are single benchmark and fractions of a point; treat them as directional, not decisive.

The drivers: why open, why this big, why now

The scale race inside China's open weight ecosystem is the first driver. Four weeks before K3, Meituan released LongCat-2.0: 1.6 trillion parameters with roughly 48 billion active per token, MIT licensed, a native 1 million token context, and agentic coding scores including 59.5 on SWE-bench Pro, trained entirely on a 50,000 card cluster of domestic Chinese accelerators.5 MarkTechPost 2026-07-05 Meituan released LongCat-2.0 on 30 June 2026 under MIT: 1.6 trillion total parameters, roughly 48 billion active per token, native 1 million token context, trained on a 50,000 card domestic Chinese accelerator cluster over 35 trillion plus tokens, SWE-bench Pro 59.5. Open source That release matched DeepSeek V4 Pro's 1.6 trillion class and proved a food delivery company could field a near frontier coder.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 5 MarkTechPost 2026-07-05 Meituan released LongCat-2.0 on 30 June 2026 under MIT: 1.6 trillion total parameters, roughly 48 billion active per token, native 1 million token context, trained on a 50,000 card domestic Chinese accelerator cluster over 35 trillion plus tokens, SWE-bench Pro 59.5. Open source Moonshot's answer was to nearly double the parameter count and claim the largest ever title outright, a marketing position no benchmark dispute can dilute.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 4 Tom's Hardware 2026-07-27 Frames Kimi K3 as the largest open weight AI model ever delivered, beating Claude Fable 5 on the Frontend Code Arena benchmark, as China works around US compute limits. Open source

The second driver is the export control environment. Tom's Hardware frames K3 explicitly as China working around US compute limits.4 Tom's Hardware 2026-07-27 Frames Kimi K3 as the largest open weight AI model ever delivered, beating Claude Fable 5 on the Frontend Code Arena benchmark, as China works around US compute limits. Open source We assess with moderate confidence that open weights are partly a distribution strategy born of constraint: a lab that cannot easily sell hosted inference into Western enterprises at scale can still make its model the substrate Western platforms sell for it. That is already happening. Together AI and Modal offered day zero hosting, with OpenRouter, Fireworks and others following shortly after, meaning US infrastructure companies began earning revenue on K3 within hours of the weights landing.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source Demand on Moonshot's own hosted side was strong enough that the company temporarily paused new subscriptions.3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source

Second order effects: who gains, who loses

The clearest winners are the neutral inference platforms. Together AI, Modal, Fireworks and OpenRouter get a frontier class product they did not have to train, priced against their GPUs rather than a lab's margin.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source Enterprises with sovereignty or confidentiality requirements gain too: a modified MIT license and downloadable weights mean teams can fine tune, quantize, or run the model fully air gapped, which no hosted API offers.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source Nvidia and the GPU cloud vendors gain by arithmetic: a model that needs 8 H100s minimum and 64 plus accelerators for recommended deployment converts open weight enthusiasm directly into hardware demand.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source 2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source

The losers are concentrated in the paid API middle tier. Closed US models that are better than K3 but not visibly so on coding and agent tasks now compete with a downloadable alternative whose marginal price is compute.3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source Smaller open model efforts, in the US and Europe, lose oxygen: when the free tier of the market is a 2.8 trillion parameter system with a 1 million token context, a 70 billion parameter release is no longer news.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source And self hosting individuals lose the plot entirely, since no consumer GPU loads even the quantized model; open weight no longer means locally runnable, and K3 makes that gap official.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source

The counter-case

The strongest argument against the scale milestone framing is that the milestone is the least useful fact about the model. Sparse activation means K3 uses 16 of 896 experts per token, so its effective compute per token is far closer to much smaller models than the 2.8 trillion headline implies, and LongCat-2.0 posts serious agentic coding scores at 48 billion active parameters.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source 5 MarkTechPost 2026-07-05 Meituan released LongCat-2.0 on 30 June 2026 under MIT: 1.6 trillion total parameters, roughly 48 billion active per token, native 1 million token context, trained on a 50,000 card domestic Chinese accelerator cluster over 35 trillion plus tokens, SWE-bench Pro 59.5. Open source If independent evaluations over the coming weeks show K3's quality per dollar of inference trailing 1.6 trillion class rivals, the largest ever label becomes a liability: an expensive to serve model whose headline number bought publicity rather than capability. The benchmark case also rests heavily on launch week leaderboards and vendor reported figures without independent replication, and the license text itself still needed confirmation in the published model card at release.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source For the thesis here to fail, hosted K3 pricing would have to settle above closed US mid tier models while its benchmark edge erodes under third party testing.

What to watch

  • Hosted K3 pricing by October 2026. If Together AI, Fireworks or OpenRouter serve K3 meaningfully below Moonshot's 3 dollar input and 15 dollar output rates, the open weights are doing their competitive job; if hosting costs keep prices at or above closed rivals, the size is a tax.1 Northflank 2026-07-17 K3 architecture and economics: 2.8 trillion total parameters, 896 experts with 16 active per token, 1,048,576 token context, MXFP4 weights with MXFP8 activations, roughly 1.4 TB raw weight download, API pricing of 3 dollars per million input and 15 dollars per million output tokens, 64 plus accelerators recommended. Open source 2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source
  • Independent benchmark replication within eight weeks. Watch whether third party evaluations confirm the Frontend Code Arena and SWE Marathon leads over US models or reduce them to noise.3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source 4 Tom's Hardware 2026-07-27 Frames Kimi K3 as the largest open weight AI model ever delivered, beating Claude Fable 5 on the Frontend Code Arena benchmark, as China works around US compute limits. Open source
  • The license text in practice. If the modified MIT terms carry attribution or revenue conditions that deter a named US enterprise deployment by year end, the openness is narrower than the headline.2 Explainx 2026-07-27 Weights published 26 July 2026 around 7:30 PM EDT ahead of the 27 July target, on Hugging Face under an expected modified MIT license; day zero hosting from Together AI and Modal with OpenRouter and Fireworks following; minimum 8 H100 80GB GPUs, no consumer GPU can load the model; K3 overtakes the 1.6T DeepSeek V4 Pro as largest tracked open weight model. Open source 3 Amplifi Labs 2026-07-27 Modified MIT license; Program Bench 77.8 vs GPT-5.6 Sol 77.6, SWE Marathon 42.0 vs Opus 4.8 40.0, Frontend Code Arena 1,679 Elo first place, GPQA 93.5 percent; trails Claude Fable 5 and GPT-5.6 Sol on general capability and deep repository analysis; Moonshot paused new subscriptions after a demand surge; tool calling and terminal use as first class capabilities. Open source
  • The next escalation from DeepSeek or Meituan. A 3 trillion plus open weight release from either within six months would confirm the scale contest is now the organizing dynamic of Chinese open AI; a pivot to smaller, cheaper agentic models would say efficiency won.5 MarkTechPost 2026-07-05 Meituan released LongCat-2.0 on 30 June 2026 under MIT: 1.6 trillion total parameters, roughly 48 billion active per token, native 1 million token context, trained on a 50,000 card domestic Chinese accelerator cluster over 35 trillion plus tokens, SWE-bench Pro 59.5. Open source
  • US policy reaction by early 2027. Any concrete US move to restrict or discourage domestic hosting of Chinese open weights would mark the moment open model diffusion became an explicit policy problem rather than a market outcome.4 Tom's Hardware 2026-07-27 Frames Kimi K3 as the largest open weight AI model ever delivered, beating Claude Fable 5 on the Frontend Code Arena benchmark, as China works around US compute limits. Open source

The durable change is not the parameter record, which will fall. It is that the reference free model for agentic coding work, the one US platforms host and US developers reach for by default, is now trained in Beijing, and each release like this one makes that default harder to unwind.