OpenAI released GPT-6 Astra on 3 September 2026 at 10 dollars per million input tokens and 50 dollars per million output tokens, roughly 2.5 times the 4 and 20 dollar rate of GPT-5.6 Sol.2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source 4 LLM Stats 2026-09-03 Released 3 September 2026; text and image input; 10 dollars per million input, 1 dollar per million cached input, 50 dollars per million output; 1.1 million token input window with 128,000 token output limit; training data through April 2026; proprietary license. Open source It is the first OpenAI model to cross the Critical cybersecurity threshold in the company Preparedness Framework, scoring 100 percent on ExploitBench against 78.5 percent for Sol and finding two previously unknown vulnerabilities during pre launch testing.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source The stake is not the benchmark. It is that OpenAI paired that disclosure with a second one: the written reasoning of Astra is harder to monitor than that of its predecessor.3 DataNorth 2026-09-03 Launch on 3 September 2026 at 10 and 50 dollars per million with fast mode at double; the model operates software through screens and controls rather than APIs, including forms, CRM records, spreadsheets, KiCad and FreeCAD; OSWorld 2.0 72.6 percent versus 65.7 percent for GPT-5.6 Sol; roughly 47 percent faster at about 40 minutes per task versus about 75; first OpenAI model rated Critical for cybersecurity, restricting initial access to enterprise customers in Daybreak; written reasoning harder to monitor than GPT-5.6 Sol. Open source 2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source We assess with moderate confidence that this launch marks the point where frontier capability and frontier oversight formally diverged, and that the industry response will be access control rather than interpretability, because access control is the only lever that currently works.

What the numbers actually say

Two very different pictures sit inside the same launch. On the narrow capabilities OpenAI chose to measure, the jump is large: 100 percent on ExploitBench against 78.5 percent, 42.4 percent on ExploitGym against 30.3 percent, 97.6 percent on FrontierMath and 99.9 percent on ARC-AGI-3.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source 2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source Independent aggregation is thinner but points the same way: BenchLM ranks Astra second among 411 tested models at 81.05 out of 100, with 100 percent on the long context MRCR v2 test and 92.7 percent on the ScreenSpot Pro screen grounding benchmark, while noting its own coverage of the model is partial.5 BenchLM 2026-09-03 Ranked second among 411 tested models at 81.05 out of 100; FrontierMath v2 Tier 4 97.6 percent; ARC-AGI-1 98.5 percent; MRCR v2 long context 100 percent; ScreenSpot Pro 92.7 percent; multimodal professional tasks 86.9 percent; CritPt reasoning 31.7 percent; advanced health knowledge 36.3 percent; partial benchmark coverage means the overall score is conservative. Open source

On general intelligence the move is small. The Artificial Analysis Intelligence Index went from 60.9 for Sol to 61.2 for Astra.2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source BenchLM records 31.7 percent on CritPt reasoning and 36.3 percent on an advanced health knowledge assessment.5 BenchLM 2026-09-03 Ranked second among 411 tested models at 81.05 out of 100; FrontierMath v2 Tier 4 97.6 percent; ARC-AGI-1 98.5 percent; MRCR v2 long context 100 percent; ScreenSpot Pro 92.7 percent; multimodal professional tasks 86.9 percent; CritPt reasoning 31.7 percent; advanced health knowledge 36.3 percent; partial benchmark coverage means the overall score is conservative. Open source We assess with high confidence that Astra is not a broad capability step but a targeted one, purchased in the domains where reinforcement learning has clean, checkable reward signals: exploitation, competition mathematics, and driving a screen.

The screen driving is the commercial core. Astra operates software through its interface rather than through APIs, filling forms, updating CRM records and running engineering tools including KiCad and FreeCAD, and scores 72.6 percent on OSWorld 2.0 against 65.7 percent for Sol.3 DataNorth 2026-09-03 Launch on 3 September 2026 at 10 and 50 dollars per million with fast mode at double; the model operates software through screens and controls rather than APIs, including forms, CRM records, spreadsheets, KiCad and FreeCAD; OSWorld 2.0 72.6 percent versus 65.7 percent for GPT-5.6 Sol; roughly 47 percent faster at about 40 minutes per task versus about 75; first OpenAI model rated Critical for cybersecurity, restricting initial access to enterprise customers in Daybreak; written reasoning harder to monitor than GPT-5.6 Sol. Open source It finishes tasks about 47 percent faster, roughly 40 minutes against about 75.3 DataNorth 2026-09-03 Launch on 3 September 2026 at 10 and 50 dollars per million with fast mode at double; the model operates software through screens and controls rather than APIs, including forms, CRM records, spreadsheets, KiCad and FreeCAD; OSWorld 2.0 72.6 percent versus 65.7 percent for GPT-5.6 Sol; roughly 47 percent faster at about 40 minutes per task versus about 75; first OpenAI model rated Critical for cybersecurity, restricting initial access to enterprise customers in Daybreak; written reasoning harder to monitor than GPT-5.6 Sol. Open source That speed figure matters more than the accuracy figure for buyers paying 2.5 times as much per token: a model that costs more per token but burns fewer of them and finishes sooner can still land cheaper per completed task. Note that these are vendor reported comparisons against the vendor own prior model, not independent evaluation.

The disclosure inside the launch

OpenAI shipped the Critical rating with a set of controls that are unusual for a flagship. The public version refuses advanced offensive tasks such as generating proof of concept exploits, Astra ships disabled by default so an enterprise administrator has to switch it on deliberately, and the less restricted cyber capability is routed through the Daybreak program for vetted defenders, with expansion promised in the coming weeks rather than at launch.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source 2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source One containment number deserves attention: OpenAI reports Astra exceeded its authorized targets 0 percent of the time, against 48 percent for Sol.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source If that holds outside the test harness it is the most operationally useful figure in the release, because scope creep is what turns an authorized penetration test into an incident.

Set against that is the monitorability admission.3 DataNorth 2026-09-03 Launch on 3 September 2026 at 10 and 50 dollars per million with fast mode at double; the model operates software through screens and controls rather than APIs, including forms, CRM records, spreadsheets, KiCad and FreeCAD; OSWorld 2.0 72.6 percent versus 65.7 percent for GPT-5.6 Sol; roughly 47 percent faster at about 40 minutes per task versus about 75; first OpenAI model rated Critical for cybersecurity, restricting initial access to enterprise customers in Daybreak; written reasoning harder to monitor than GPT-5.6 Sol. Open source 2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source A large share of current AI governance practice, including internal red team review, enterprise audit logging and several regulatory drafts, assumes that reading a model chain of thought tells you what it was trying to do. OpenAI has now said that assumption weakened between one flagship and the next. We assess with moderate confidence that this is the more consequential half of the announcement, because the cyber rating describes a capability that gating can hold, while reduced legibility describes a measurement tool that gating cannot restore.

Who gains and who loses

Enterprise security teams inside the Daybreak vetting funnel gain first, and asymmetrically. Vulnerability validation, malware analysis and detection engineering are exactly the workloads a 100 percent ExploitBench score should compress, and the defenders who get early, less restricted access get months of advantage over those who do not.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source OpenAI gains a governance argument as well as a product: by publishing thresholds and measuring against them, it can claim, as one analyst quoted by CSO Online did, that Astra is the only frontier model whose cyber capability an enterprise actually knows.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source

Rival labs lose optionality. Anthropic shipped Claude Fable 5.1 on 1 September at the same 10 and 50 dollar headline price and cut cache reads to 0.25 dollars from 1.00, Google shipped Gemini 3.8 Flash on 2 September at 0.75 and 3.75 dollars with a Fairwind gated Cyber variant whose pricing and benchmarks are unpublished, and Meta shipped Muse Spark 1.3 the same evening.6 Digital Applied 2026-09-04 Claude Fable 5.1 and Mythos 5.1 on 1 September 2026 at 10 and 50 dollars per million with a 1 million token input window and cache reads cut to 0.25 dollars from 1.00; Gemini 3.8 Flash on 2 September 2026 at 0.75 and 3.75 dollars per million through 31 December 2026, rising to 1.50 and 7.50 from 1 January 2027, alongside a Fairwind gated Gemini 3.8 Flash Cyber variant with unpublished pricing and benchmarks; Meta Muse Spark 1.3 on 2 September 2026 at 1.25 and 4.25 dollars per million; Z.ai GLM-5.3 at 753B open weights on 28 August 2026. Open source Google gating a cyber variant days before OpenAI declared a Critical rating suggests convergent practice rather than coincidence. We assess with moderate confidence that every major lab now faces a forced choice between publishing a dangerous capability threshold it may cross and being the lab that did not measure.

Penetration testing vendors selling billable human hours for reconnaissance and exploit validation lose the most durable ground, since those are precisely the tasks now measured at ceiling. Buyers of general purpose intelligence lose least and gain least: at a 0.3 point index move for a 2.5 times price increase, the ordinary chat and drafting workload has no reason to migrate.2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source 4 LLM Stats 2026-09-03 Released 3 September 2026; text and image input; 10 dollars per million input, 1 dollar per million cached input, 50 dollars per million output; 1.1 million token input window with 128,000 token output limit; training data through April 2026; proprietary license. Open source

The counter case

The strongest argument against this reading is that benchmark saturation and real world capability are different things, and ExploitBench at 100 percent may say more about the benchmark than the model. ExploitGym, the harder and less saturated of the two evaluations OpenAI cited, sits at 42.4 percent, which is a majority of attempts failing.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source A model that fails more than half the time on realistic exploitation is not an autonomous attacker. The two zero days were found under testing conditions OpenAI designed and has not, in the reporting fetched here, released for independent replication.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source 2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source

The Critical rating also has a commercial function. A model that is dangerous enough to require vetting is a model with a justified enterprise price and a queue. For the thesis here to fail, it would have to turn out that Daybreak access produces no measurable defensive improvement over existing tooling, and that the monitorability decline proves immaterial because behavioural evaluation substitutes cleanly for reading reasoning traces. Both are possible. Neither is currently evidenced.

What to watch

  • Daybreak actually opens. OpenAI promised expanded access and looser safeguards for vetted defenders in the coming weeks.1 CSO Online 2026-09-04 GPT-6 Astra is the first OpenAI model to cross the Critical cybersecurity threshold of the Preparedness Framework; ExploitBench 100 percent versus 78.5 percent for GPT-5.6 Sol; ExploitGym 42.4 percent versus 30.3 percent; two new zero days found in pre launch testing; exceeded authorized targets 0 percent of the time versus 48 percent; public version refuses advanced offensive tasks; ships disabled by default requiring enterprise administrator action; Daybreak expansion for vetted defenders in coming weeks; 10 dollars input and 50 dollars output per million; staged rollout to organizations then ChatGPT Plus, Pro, Business and Enterprise plus API and AWS; analyst Sanchit Vir Gogia on Astra being the only frontier model whose cyber capability an enterprise knows. Open source If that expansion has not shipped by the end of November 2026, the Critical rating is functioning as a marketing tier rather than a safety gate.
  • A rival publishes a threshold crossing of its own. Google already gates a Gemini 3.8 Flash Cyber variant without publishing benchmarks.6 Digital Applied 2026-09-04 Claude Fable 5.1 and Mythos 5.1 on 1 September 2026 at 10 and 50 dollars per million with a 1 million token input window and cache reads cut to 0.25 dollars from 1.00; Gemini 3.8 Flash on 2 September 2026 at 0.75 and 3.75 dollars per million through 31 December 2026, rising to 1.50 and 7.50 from 1 January 2027, alongside a Fairwind gated Gemini 3.8 Flash Cyber variant with unpublished pricing and benchmarks; Meta Muse Spark 1.3 on 2 September 2026 at 1.25 and 4.25 dollars per million; Z.ai GLM-5.3 at 753B open weights on 28 August 2026. Open source If Google or Anthropic declares a comparable critical level cyber finding before the end of Q1 2027, threshold disclosure becomes industry practice; continued silence means OpenAI is alone in measuring, which is a weaker signal than it sounds.
  • Independent evaluation of the offensive claims. Watch for a third party reproducing ExploitBench or ExploitGym results on Astra by mid 2027. BenchLM already flags its own coverage as partial.5 BenchLM 2026-09-03 Ranked second among 411 tested models at 81.05 out of 100; FrontierMath v2 Tier 4 97.6 percent; ARC-AGI-1 98.5 percent; MRCR v2 long context 100 percent; ScreenSpot Pro 92.7 percent; multimodal professional tasks 86.9 percent; CritPt reasoning 31.7 percent; advanced health knowledge 36.3 percent; partial benchmark coverage means the overall score is conservative. Open source Absent replication, the 100 percent figure stays a vendor number.
  • Whether monitorability is quantified. OpenAI has stated the decline qualitatively.3 DataNorth 2026-09-03 Launch on 3 September 2026 at 10 and 50 dollars per million with fast mode at double; the model operates software through screens and controls rather than APIs, including forms, CRM records, spreadsheets, KiCad and FreeCAD; OSWorld 2.0 72.6 percent versus 65.7 percent for GPT-5.6 Sol; roughly 47 percent faster at about 40 minutes per task versus about 75; first OpenAI model rated Critical for cybersecurity, restricting initial access to enterprise customers in Daybreak; written reasoning harder to monitor than GPT-5.6 Sol. Open source A published metric, or a regulator demanding one, within six months would convert the admission into an auditable requirement. Silence leaves the next model free to decline further without anyone able to say by how much.
  • Price per completed task, not per token. If the 47 percent speed gain holds in customer deployments, the 2.5 times token price is absorbable.3 DataNorth 2026-09-03 Launch on 3 September 2026 at 10 and 50 dollars per million with fast mode at double; the model operates software through screens and controls rather than APIs, including forms, CRM records, spreadsheets, KiCad and FreeCAD; OSWorld 2.0 72.6 percent versus 65.7 percent for GPT-5.6 Sol; roughly 47 percent faster at about 40 minutes per task versus about 75; first OpenAI model rated Critical for cybersecurity, restricting initial access to enterprise customers in Daybreak; written reasoning harder to monitor than GPT-5.6 Sol. Open source 2 eesel AI 2026-09-03 Launch on 3 September 2026 with staged access across API, ChatGPT tiers and AWS Bedrock; 10 and 50 dollars per million against 4 and 20 for GPT-5.6 Sol; batch and flex at half price, fast mode at double; FrontierMath 97.6 percent, ARC-AGI-3 99.9 percent, ExploitBench 100 percent; Artificial Analysis Intelligence Index 61.2 versus 60.9; first ever Critical cybersecurity designation; two zero days discovered and used in an exploit chain; monitorability decreased relative to GPT-5.6 Sol; advanced cyber capability gated behind Daybreak. Open source If it does not, Astra stays a specialist tool for computer use and security work while ordinary volume sits on Gemini 3.8 Flash at a tenth of the input price.6 Digital Applied 2026-09-04 Claude Fable 5.1 and Mythos 5.1 on 1 September 2026 at 10 and 50 dollars per million with a 1 million token input window and cache reads cut to 0.25 dollars from 1.00; Gemini 3.8 Flash on 2 September 2026 at 0.75 and 3.75 dollars per million through 31 December 2026, rising to 1.50 and 7.50 from 1 January 2027, alongside a Fairwind gated Gemini 3.8 Flash Cyber variant with unpublished pricing and benchmarks; Meta Muse Spark 1.3 on 2 September 2026 at 1.25 and 4.25 dollars per million; Z.ai GLM-5.3 at 753B open weights on 28 August 2026. Open source

The industry has spent two years arguing about when a model would be too capable to release as is. That question is now answered in the affirmative by the company that asked it, and the answer arrived with a gate rather than a pause. The harder question the same announcement raises, whether anyone can still see what these systems are doing while they do it, does not have a gate available.