The Institute of Foundation Models at Mohamed bin Zayed University of Artificial Intelligence in Abu Dhabi released K2 Horizon on 3 September: six models from 0.9 billion to 375 billion parameters, all under Apache 2.0, all described as fully open, meaning weights, code, training data and methodology.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source The flagship is a 375 billion parameter mixture of experts with 23 billion active, which Artificial Analysis scores at 38 on its Intelligence Index, eleventh of 112 in its category and well above the open weight average.3 Artificial Analysis 2026-09-03 Intelligence Index 38; ranked 11 of 112 in category; 524k context; 375B total, 23B active; no speed or price data. Open source The claim of full openness is true today for the 3.7 billion and 7 billion models and is a promise for the two largest, whose training data and code the model cards say will be released.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source Our assessment, with high confidence, is that this is the most complete open release at this scale from any institution outside China; with moderate confidence, that the 7 billion model is the most consequential piece; and with moderate confidence that the gap between released and promised is a schedule, not a hedge, though it should be watched.
What shipped
The six sizes are 0.9B, 3.7B, 7B, 32B dense, 36B with 4B active and 375B with 23B active.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source The institute describes three architectural features: diffusion distillation that generates blocks of tokens in parallel for about a threefold speedup, a Mixture of Value Attention it says improves reasoning without added compute, and dynamic routing that sends a task to the cheapest model in the fleet that can do it.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source Training used about 20 trillion tokens per model, with 17 percent explicit reasoning trajectories and about 10 trillion synthetic tokens, per CellCog's reading of the documentation.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source Distribution is immediate: Hugging Face, vLLM, SGLang and Ollama on day one, with hosted inference from Compass, Cerebras, AWS and Nebius.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source 2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source
The benchmark that stands out is the small one. The 7B scored 70.6 percent on SWE-bench Verified against 50.8 for Qwen3.5-9B, a twenty point lead at a size that runs on a laptop.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source The flagship's numbers are mixed: 65.3 percent on Toolathlon Verified against 59.9 for GLM 5.2, but 66.9 percent on Terminal-Bench 2.1 after the institute's own reward hacking audit removed 24 flagged trials, against 77.9 for GLM 5.2.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source That the institute audited its own result downward and published the correction is itself a data point about how it intends to operate.
Fully open, in stages
The distinction CellCog draws is the one that matters: open weights means the model and the recipe to run it; fully open means the training lifecycle, data included, so the result can be reproduced and improved rather than merely used.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source K2 Horizon delivers the latter for the two mid sized models now, with intermediate checkpoints, and states the rest will follow.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source Compute costs are not disclosed anywhere, which is the one gap the institute has not promised to close.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source
Professor Eric Xing's framing, that progress depends on the ability to examine and build on the technology rather than access it through an API, is the argument for releasing data.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source The staging is understandable: a 375 billion parameter training corpus is a very large artifact to clear for release. It is also the part that competitors would most like to see, and the part most likely to slip.
Who gains and who loses
Developers who need a strong small coding model gain the most, immediately, if the 7B result holds in independent testing.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source Researchers gain a reproducible mid scale pipeline with data, which almost no lab provides. Abu Dhabi gains a credible claim to leadership in open AI at a moment when the largest open models otherwise come from China.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source The inference partners, Cerebras among them, gain a flagship open model to serve.
The Chinese open weight labs lose a little of their monopoly on frontier scale open releases, though GLM 5.2 still leads the flagship on the hardest agentic benchmark.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source Meta, whose open strategy has been quiet this year, loses ground on the openness axis to an institution that publishes training data. And the open weight but closed data labs everywhere lose the argument that full openness at scale is impractical, if the 375B data ships.
The counter case
The assessment that the staging is a schedule rather than a hedge could be wrong. Institutions have announced full releases before and delivered weights only, and the incentive to withhold the largest training corpus grows once the model is being used. If the 375B data and code do not appear, K2 Horizon is an open weight release with unusually good documentation, which is valuable but not what was claimed.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source The benchmarks are also the institute's own, with the corrected Terminal-Bench figure a reminder that agentic benchmarks are easy to inflate and that the flagship trails the best open model on that test.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source And an Intelligence Index of 38, eleventh in category, places the 375B as a strong open model rather than a frontier one.3 Artificial Analysis 2026-09-03 Intelligence Index 38; ranked 11 of 112 in category; 524k context; 375B total, 23B active; no speed or price data. Open source The importance of K2 Horizon rests on the openness, and the openness rests on a promise.
What to watch
- The 375B training data lands. Public release of the flagship's training data and code by the end of October would make the fully open claim true in full; a slip past the end of 2026 would make it a weights release.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source
- Independent 7B coding results. A third party reproducing a SWE-bench Verified score near 70 percent for the 7B within two months would confirm the most useful result in the release.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source
- Derivative models appear. Fine tunes or continued pre trains built on the released data, not only the weights, within six months would show the full openness is being used as intended.1 MBZUAI 2026-09-03 Six sizes from 0.9B to 375B-A23B; Apache 2.0 with weights, code, data and methodology; diffusion distillation, Mixture of Value Attention, dynamic routing; 7B claimed best under 10B; Hugging Face, vLLM, SGLang plus Compass, Cerebras, AWS and Nebius; Xing and Liu quotes. Open source
- Compute disclosure. A published training compute figure for the fleet would close the last transparency gap; continued silence keeps the reproducibility claim partial.2 CellCog 2026-09-03 Only the 3.7B and 7B are fully released today; 375B and 36B data and code promised; 32B stage one only; 7B SWE-bench Verified 70.6 vs 50.8 for Qwen3.5-9B; 375B Terminal-Bench 66.9 after audit vs 77.9 for GLM 5.2, Toolathlon 65.3 vs 59.9; about 20 trillion tokens, 17 percent reasoning, 10 trillion synthetic; compute undisclosed. Open source
Six models under Apache 2.0 is a fact. Fully open is, for now, a fact for two of them and a date for the rest.