Z.ai announced GLM-5.3 on 14 August 2026 and said that API access and open weights would arrive in stages after safety evaluation, a break from the pattern that put GLM-5.2 on Hugging Face within days.4 ExplainX 2026-08-14 Z.ai announced GLM-5.3 on 14 August 2026 post-trained on a 743 billion parameter base; CyberGym 84.5 percent, AutomationBench 48.2 percent, GDPVal-AA v2 1769 Elo; coding plan tiers at 12.6, 56 and 117.6 dollars a month; Z.ai said API access and open weights will be released in stages following rigorous safety evaluations, a departure from GLM-5.2; an 28 August Hugging Face target was missed with no confirmed new timeline. Open source The promised window pointed at roughly 28 August, and that date passed with no weights repository published.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source 5 ExplainX 2026-08-26 GLM-5.3-Flash launched 26 August 2026: 320 billion parameters with 18 billion active, MIT weights on Hugging Face, 1 million token multimodal context, API at 0.15 input, 0.50 output and 0.03 cached per million tokens; DeepSWE v1.1 63.4, AutomationBench 48.8, GDPVal-AA v2 1773, Terminal Bench 2.1 84.3 behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4; runs entirely on Chinese AI chips; the larger GLM-5.3 missed its 28 August weights target with release staged pending safety review. Open source The stake is that the informal floor under open weight availability, the assumption that Chinese labs will publish whatever they build because publishing is their distribution strategy, has now been tested by one of the labs that set it. We assess with moderate confidence that the stated reason, offensive security capability that grew faster than the lab expected, is at least partly genuine, and with high confidence that the operative outcome regardless of motive is a two track release model: capable but bounded models ship under permissive licenses, and the frontier line does not.
What Z.ai says it built
GLM-5.3 is not a new pretrain. It sits on the same 744 billion parameter mixture of experts base as GLM-5.2, roughly 40 billion parameters active per token, with 1 million tokens of native context, and every gain comes from extended post-training.2 Fello AI 2026-08-26 744 billion parameter base identical to GLM-5.2 with no new pretrain; Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9; CyberGym 84.5 percent and ExploitBench 54.4 percent; Z.ai claims a CyberGym lead over Claude Mythos 5 and GPT-5.6 Sol; 2,436 findings across 269 projects with 1,097 critical or high; GLM Coding Plan at 18 to 160 dollars a month; as of the 26 August update no GLM-5.3 weights repository on Hugging Face. Open source 3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source Part of that post-training used data and environments built specifically for finding software vulnerabilities.1 The Decoder 2026-08-14 GLM-5.3 released 14 August 2026 with gains from extended post-training on the GLM-5.2 base; trained on vulnerability discovery data and environments; Z.ai documentation describes reasoning across multiple stages of exploitation and forming complete exploitation chains; 2,436 vulnerabilities across 269 projects; available via the GLM Coding Plan with open weights roughly two weeks out pending security reviews. Open source The self reported movement on security benchmarks is the largest in the release: CyberGym from 77.2 to 84.5, ExploitBench from 24.4 to 54.4, and ExploitGym from 29 and 39 solves at two and six hours to 105 and 130.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source Coding numbers moved too, with Terminal-Bench 3.0 going from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9 against the prior model.2 Fello AI 2026-08-26 744 billion parameter base identical to GLM-5.2 with no new pretrain; Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9; CyberGym 84.5 percent and ExploitBench 54.4 percent; Z.ai claims a CyberGym lead over Claude Mythos 5 and GPT-5.6 Sol; 2,436 findings across 269 projects with 1,097 critical or high; GLM Coding Plan at 18 to 160 dollars a month; as of the 26 August update no GLM-5.3 weights repository on Hugging Face. Open source
The lab is unusually direct about what surprised it. Its own documentation describes the model beginning to reason "across multiple stages of exploitation", assembling plans for full exploitation chains rather than isolated bug finds.1 The Decoder 2026-08-14 GLM-5.3 released 14 August 2026 with gains from extended post-training on the GLM-5.2 base; trained on vulnerability discovery data and environments; Z.ai documentation describes reasoning across multiple stages of exploitation and forming complete exploitation chains; 2,436 vulnerabilities across 269 projects; available via the GLM Coding Plan with open weights roughly two weeks out pending security reviews. Open source Alongside the model, Z.ai published a ledger: 2,436 vulnerabilities found across 269 open source projects, of which 1,097 are rated critical or high.1 The Decoder 2026-08-14 GLM-5.3 released 14 August 2026 with gains from extended post-training on the GLM-5.2 base; trained on vulnerability discovery data and environments; Z.ai documentation describes reasoning across multiple stages of exploitation and forming complete exploitation chains; 2,436 vulnerabilities across 269 projects; available via the GLM Coding Plan with open weights roughly two weeks out pending security reviews. Open source 2 Fello AI 2026-08-26 744 billion parameter base identical to GLM-5.2 with no new pretrain; Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9; CyberGym 84.5 percent and ExploitBench 54.4 percent; Z.ai claims a CyberGym lead over Claude Mythos 5 and GPT-5.6 Sol; 2,436 findings across 269 projects with 1,097 critical or high; GLM Coding Plan at 18 to 160 dollars a month; as of the 26 August update no GLM-5.3 weights repository on Hugging Face. Open source The breakdown reported is 107 critical, 990 high, 1,286 medium and 53 low, with only 53 findings publicly disclosed and 2,383 held under embargo, an average of 26.6 years undiscovered and the oldest defect dating to 1981.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source Note the sourcing quality here: every benchmark and every count in this section comes from Z.ai charts and Z.ai documentation relayed by trade outlets, not from an independent evaluation, and the artifact that would let a third party check any of it is precisely the thing being withheld.
The week the two track pattern became visible
What makes the missed date legible rather than routine is what Z.ai did instead. On 26 August it launched GLM-5.3-Flash, a separate product line: 320 billion total parameters with 18 billion active, MIT licensed weights on Hugging Face, 1 million token context, text, image and video input, and API pricing at 0.15 dollars per million input tokens against 0.50 for output.5 ExplainX 2026-08-26 GLM-5.3-Flash launched 26 August 2026: 320 billion parameters with 18 billion active, MIT weights on Hugging Face, 1 million token multimodal context, API at 0.15 input, 0.50 output and 0.03 cached per million tokens; DeepSWE v1.1 63.4, AutomationBench 48.8, GDPVal-AA v2 1773, Terminal Bench 2.1 84.3 behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4; runs entirely on Chinese AI chips; the larger GLM-5.3 missed its 28 August weights target with release staged pending safety review. Open source Flash is not weak. It posts 63.4 on DeepSWE v1.1 and 84.3 on Terminal Bench 2.1, close behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4, and Z.ai says it runs entirely on Chinese AI chips.5 ExplainX 2026-08-26 GLM-5.3-Flash launched 26 August 2026: 320 billion parameters with 18 billion active, MIT weights on Hugging Face, 1 million token multimodal context, API at 0.15 input, 0.50 output and 0.03 cached per million tokens; DeepSWE v1.1 63.4, AutomationBench 48.8, GDPVal-AA v2 1773, Terminal Bench 2.1 84.3 behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4; runs entirely on Chinese AI chips; the larger GLM-5.3 missed its 28 August weights target with release staged pending safety review. Open source The same day, Alibaba published Qwen3.8-Flash-Next under Apache 2.0: 125 billion parameters with 6 billion active, research weights on Hugging Face and ModelScope, and a claim of better results than Qwen3.7-Plus at roughly one-ninth the training cost.6 The Decoder 2026-08-26 Qwen3.8-Flash-Next released 26 August 2026: 125 billion total parameters with 6 billion active, a 51 billion parameter N-gram embedding layer in system RAM, 262,144 native context to 1 million with YaRN, Apache 2.0 research weights on Hugging Face and ModelScope, production pricing 0.16 input and 0.47 output per million tokens; better results than Qwen3.7-Plus at roughly one-ninth the training cost; DeepSWE 58.7, SWE-bench Pro 62.5 against Claude Opus 4.6 at 53.4, CoWorkBench 73.9, JobBench 55.7. Open source
So the open weight pipeline out of China did not slow in the window. It sorted. Two labs put genuinely capable, cheap, permissively licensed models on public repositories inside 24 hours of each other, while the one model in the group trained specifically to find exploits stayed behind a paid coding plan priced by tier.4 ExplainX 2026-08-14 Z.ai announced GLM-5.3 on 14 August 2026 post-trained on a 743 billion parameter base; CyberGym 84.5 percent, AutomationBench 48.2 percent, GDPVal-AA v2 1769 Elo; coding plan tiers at 12.6, 56 and 117.6 dollars a month; Z.ai said API access and open weights will be released in stages following rigorous safety evaluations, a departure from GLM-5.2; an 28 August Hugging Face target was missed with no confirmed new timeline. Open source 5 ExplainX 2026-08-26 GLM-5.3-Flash launched 26 August 2026: 320 billion parameters with 18 billion active, MIT weights on Hugging Face, 1 million token multimodal context, API at 0.15 input, 0.50 output and 0.03 cached per million tokens; DeepSWE v1.1 63.4, AutomationBench 48.8, GDPVal-AA v2 1773, Terminal Bench 2.1 84.3 behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4; runs entirely on Chinese AI chips; the larger GLM-5.3 missed its 28 August weights target with release staged pending safety review. Open source 6 The Decoder 2026-08-26 Qwen3.8-Flash-Next released 26 August 2026: 125 billion total parameters with 6 billion active, a 51 billion parameter N-gram embedding layer in system RAM, 262,144 native context to 1 million with YaRN, Apache 2.0 research weights on Hugging Face and ModelScope, production pricing 0.16 input and 0.47 output per million tokens; better results than Qwen3.7-Plus at roughly one-ninth the training cost; DeepSWE 58.7, SWE-bench Pro 62.5 against Claude Opus 4.6 at 53.4, CoWorkBench 73.9, JobBench 55.7. Open source We assess with moderate confidence that this sorting is deliberate and will be repeated, because it lets a lab keep the reputational and adoption benefits of open release while removing the one artifact that carries the clearest misuse argument.
Who gains and who loses
Z.ai gains twice. Commercially, the delay converts what would have been free self-hosting into subscription revenue for as long as it lasts, with the coding plan priced at three tiers reported at 12.6, 56 and 117.6 dollars a month by one outlet and as a range of 18 to 160 dollars by another, an inconsistency worth flagging since neither figure comes from a price sheet we read.4 ExplainX 2026-08-14 Z.ai announced GLM-5.3 on 14 August 2026 post-trained on a 743 billion parameter base; CyberGym 84.5 percent, AutomationBench 48.2 percent, GDPVal-AA v2 1769 Elo; coding plan tiers at 12.6, 56 and 117.6 dollars a month; Z.ai said API access and open weights will be released in stages following rigorous safety evaluations, a departure from GLM-5.2; an 28 August Hugging Face target was missed with no confirmed new timeline. Open source 2 Fello AI 2026-08-26 744 billion parameter base identical to GLM-5.2 with no new pretrain; Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9; CyberGym 84.5 percent and ExploitBench 54.4 percent; Z.ai claims a CyberGym lead over Claude Mythos 5 and GPT-5.6 Sol; 2,436 findings across 269 projects with 1,097 critical or high; GLM Coding Plan at 18 to 160 dollars a month; as of the 26 August update no GLM-5.3 weights repository on Hugging Face. Open source Reputationally, it acquires a safety posture cheaply, and it does so while still shipping MIT weights for a 320 billion parameter multimodal model in the same fortnight.5 ExplainX 2026-08-26 GLM-5.3-Flash launched 26 August 2026: 320 billion parameters with 18 billion active, MIT weights on Hugging Face, 1 million token multimodal context, API at 0.15 input, 0.50 output and 0.03 cached per million tokens; DeepSWE v1.1 63.4, AutomationBench 48.8, GDPVal-AA v2 1773, Terminal Bench 2.1 84.3 behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4; runs entirely on Chinese AI chips; the larger GLM-5.3 missed its 28 August weights target with release staged pending safety review. Open source
Alibaba gains the vacuum. With the strongest advertised Chinese coding and security model unavailable to download, Qwen3.8-Flash-Next is the capable open weight release of that week, and it landed under the most permissive license in the group.6 The Decoder 2026-08-26 Qwen3.8-Flash-Next released 26 August 2026: 125 billion total parameters with 6 billion active, a 51 billion parameter N-gram embedding layer in system RAM, 262,144 native context to 1 million with YaRN, Apache 2.0 research weights on Hugging Face and ModelScope, production pricing 0.16 input and 0.47 output per million tokens; better results than Qwen3.7-Plus at roughly one-ninth the training cost; DeepSWE 58.7, SWE-bench Pro 62.5 against Claude Opus 4.6 at 53.4, CoWorkBench 73.9, JobBench 55.7. Open source Western closed labs gain an argument. The standard objection to Western release restraint is that Chinese labs will publish an equivalent within months, so restraint only forfeits share. A Chinese lab publicly withholding its top model on exploitation capability grounds weakens that objection for as long as the withholding holds.4 ExplainX 2026-08-14 Z.ai announced GLM-5.3 on 14 August 2026 post-trained on a 743 billion parameter base; CyberGym 84.5 percent, AutomationBench 48.2 percent, GDPVal-AA v2 1769 Elo; coding plan tiers at 12.6, 56 and 117.6 dollars a month; Z.ai said API access and open weights will be released in stages following rigorous safety evaluations, a departure from GLM-5.2; an 28 August Hugging Face target was missed with no confirmed new timeline. Open source
Self-hosters lose most directly. Anyone who planned a local GLM-5.3 deployment has no artifact, no license text, and no quantization guidance, since Z.ai has not stated a license for 5.3 and the MIT precedent from 5.1 and 5.2 is a pattern rather than a commitment.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source Independent safety researchers lose too: the security claims that justify the withholding cannot be reproduced by anyone outside Z.ai without the weights. And the maintainers of those 269 projects sit on the wrong side of an embargo covering 2,383 findings they cannot see, on a disclosure schedule set by the finder.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source
The counter-case
The strongest honest reading against the safety thesis is that this is a two week engineering and commercial delay wearing a safety label. Nothing in the reporting establishes that Z.ai committed to 28 August as a hard date; the figure comes from a two week statement at launch and a placeholder on a repository page, and one outlet notes the company did not guarantee it.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source 4 ExplainX 2026-08-14 Z.ai announced GLM-5.3 on 14 August 2026 post-trained on a 743 billion parameter base; CyberGym 84.5 percent, AutomationBench 48.2 percent, GDPVal-AA v2 1769 Elo; coding plan tiers at 12.6, 56 and 117.6 dollars a month; Z.ai said API access and open weights will be released in stages following rigorous safety evaluations, a departure from GLM-5.2; an 28 August Hugging Face target was missed with no confirmed new timeline. Open source Large weight drops slip for mundane reasons. Monetizing API access before free self-hosting is available is an obvious motive that requires no safety story at all, and the Flash launch on 26 August shows the release pipeline is running normally.5 ExplainX 2026-08-26 GLM-5.3-Flash launched 26 August 2026: 320 billion parameters with 18 billion active, MIT weights on Hugging Face, 1 million token multimodal context, API at 0.15 input, 0.50 output and 0.03 cached per million tokens; DeepSWE v1.1 63.4, AutomationBench 48.8, GDPVal-AA v2 1773, Terminal Bench 2.1 84.3 behind Opus 4.8 at 85.0 and GPT-5.6 Terra at 87.4; runs entirely on Chinese AI chips; the larger GLM-5.3 missed its 28 August weights target with release staged pending safety review. Open source
For our thesis to fail, the GLM-5.3 weights should appear within weeks under MIT, materially unchanged, with no published safety evaluation and no capability restrictions. That outcome would mark this as a schedule slip that a trade press cycle read as policy. For the thesis to hold, the weights either stay unreleased, arrive under a more restrictive license, or arrive with the security capability visibly reduced. Both paths are live, and the evidence available today is entirely Z.ai self reporting relayed through four trade outlets, none of which conducted an independent evaluation.
What to watch
- The weights, and the license on them. If a GLM-5.3 repository appears by the end of September 2026 under MIT with no stated capability changes, read this as a slip.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source A restrictive license, a filtered checkpoint, or silence past October makes it policy.
- Whether the next Zhipu frontier model repeats the pattern. One staged release is an exception. If the following GLM flagship also launches API first with weights staged behind evaluation, the two track model is the lab's standing policy.4 ExplainX 2026-08-14 Z.ai announced GLM-5.3 on 14 August 2026 post-trained on a 743 billion parameter base; CyberGym 84.5 percent, AutomationBench 48.2 percent, GDPVal-AA v2 1769 Elo; coding plan tiers at 12.6, 56 and 117.6 dollars a month; Z.ai said API access and open weights will be released in stages following rigorous safety evaluations, a departure from GLM-5.2; an 28 August Hugging Face target was missed with no confirmed new timeline. Open source
- The embargo clock on 2,383 findings. Watch whether disclosures move beyond the 53 already public, and on what schedule.3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source A ledger that stays sealed into 2027 turns a security credential into an unverifiable claim.
- Whether Alibaba, DeepSeek or Moonshot adopt staged release language. Qwen shipped Apache 2.0 research weights on the same day Z.ai held its own back.6 The Decoder 2026-08-26 Qwen3.8-Flash-Next released 26 August 2026: 125 billion total parameters with 6 billion active, a 51 billion parameter N-gram embedding layer in system RAM, 262,144 native context to 1 million with YaRN, Apache 2.0 research weights on Hugging Face and ModelScope, production pricing 0.16 input and 0.47 output per million tokens; better results than Qwen3.7-Plus at roughly one-ninth the training cost; DeepSWE 58.7, SWE-bench Pro 62.5 against Claude Opus 4.6 at 53.4, CoWorkBench 73.9, JobBench 55.7. Open source If any of the other three attaches an evaluation gate to a flagship release before the end of 2026, the norm has moved, not just one lab.
- Independent reproduction of the cyber numbers. CyberGym at 84.5 and ExploitBench at 54.4 are Z.ai figures.2 Fello AI 2026-08-26 744 billion parameter base identical to GLM-5.2 with no new pretrain; Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9; CyberGym 84.5 percent and ExploitBench 54.4 percent; Z.ai claims a CyberGym lead over Claude Mythos 5 and GPT-5.6 Sol; 2,436 findings across 269 projects with 1,097 critical or high; GLM Coding Plan at 18 to 160 dollars a month; as of the 26 August update no GLM-5.3 weights repository on Hugging Face. Open source 3 Modem Guides 2026-08-14 Open weights promised roughly two weeks after the 14 August launch, around 28 August; 744 billion parameter mixture of experts with about 40 billion active and 1 million token native context; no stated license for 5.3, with MIT on 5.1 and 5.2 described as a pattern not a promise; ledger of 2,436 findings across 269 projects, 107 critical, 990 high, 1,286 medium, 53 low, 53 disclosed and 2,383 embargoed, 26.6 year average and oldest flaw from 1981; CyberGym 77.2 to 84.5, ExploitBench 24.4 to 54.4, ExploitGym 29/39 to 105/130. Open source Confirmation by an outside evaluator, whenever the artifact allows it, is what would turn the justification for withholding into evidence.
The interesting fact is not that one release date slipped. It is that the argument Western labs have used to justify holding capable weights back now has a Chinese precedent attached to it, offered voluntarily by a lab whose entire distribution strategy was giving the weights away.