Google released Gemini 3.8 Flash on 2 September, its fourth Flash model in under four months, generally available through the Gemini API, AI Studio and Android Studio.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source 2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source On Google's own numbers it ties Claude Opus 5 on DeepSWE v1.1 and Terminal-bench 2.1 and edges it on HLE-Verified and the Vals finance agent benchmark, at 0.75 dollars per million input tokens and 3.75 per million output against Opus 5's 25 dollar output price.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source Two facts complicate the bargain. Google says the model is built on 3.7 Flash rather than a new base and works harder by spending more thinking tokens, which Artificial Analysis measures as a 30 percent rise in output tokens per task and a 40 percent rise in cost per task despite unchanged per token pricing.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source And the price is a promotion: on 1 January 2027 it doubles to 1.50 and 7.50.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source Our assessment, with high confidence, is that 3.8 Flash is a real capability gain at the coding and agent tasks Google chose to publish; with moderate confidence, that the effective price for a given task will roughly triple between now and February once the token inflation and the schedule compound; and with moderate confidence that Google's own advice to stay on 3.7 Flash for cost sensitive work is the honest summary.

What the benchmarks show

Against its predecessor the gains are broad. DeepSWE v1.1 rose to 73.7 percent from 65.3, OSWorld-2.0 to 59.0 from 50.6, Terminal-bench 2.1 to 89.4 from 85.8 and Terminal-bench 4.0 to 19.1 from 11.2.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source DataCamp's account of Google's table puts Terminal-Bench 2.1 at 90.8 against 81.6 and adds SWE-Bench Pro at 61.6 and SWE-Atlas at 51.9; the two Terminal-bench figures differ between outlets, and this brief treats the discrepancy as a reporting artifact rather than a fact.4 DataCamp 2026-09-02 Terminal-Bench 2.1 at 90.8 vs 81.6 in this account; SWE-Bench Pro 61.6; SWE-Atlas 51.9; HLE flat at 45.4 vs 45.7; Flash Cyber vulnerability discovery above 70 percent, CWE-Bench 47.2; Fairwind gating; defensive bias. Open source Humanity's Last Exam was flat at 45.4 against 45.7.4 DataCamp 2026-09-02 Terminal-Bench 2.1 at 90.8 vs 81.6 in this account; SWE-Bench Pro 61.6; SWE-Atlas 51.9; HLE flat at 45.4 vs 45.7; Flash Cyber vulnerability discovery above 70 percent, CWE-Bench 47.2; Fairwind gating; defensive bias. Open source

Against Opus 5 the picture is mixed in a specific way. Ties on DeepSWE at 73.7 versus 74.0 and Terminal-bench 2.1 at 89.4 versus 89.1; a lead on Vals Finance Agent v2 at 61.4 versus 58.6 and on Harvey's Legal Agent at 10.0 versus 6.7; a large deficit on OSWorld-2.0 computer use at 59.0 versus 75.4.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source Flash matches the frontier where the task is text and code in a terminal, and trails where the task is operating a screen.

Independent measurement agrees on direction. Artificial Analysis scores the model 59 on its Intelligence Index, up 3 from 3.7 Flash, level with GPT-5.6 Sol at sub maximum reasoning and with Grok 4.6, and 17th of 196 models.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source It is third fastest on output at about 300 tokens per second but slow to start, with a 13.30 second time to first token against a 2.99 second median.3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source

The catch, in numbers

The model reasons more. Output tokens per task rose about 30 percent to 48,000, verbosity on the Artificial Analysis suite is 120 million tokens against a 71 million median, and time per task at high reasoning rose from 2.2 to 2.5 minutes.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source Because thinking tokens bill as output, cost per task rose about 40 percent to 0.58 dollars at high reasoning, 0.41 at medium and 0.24 at low, with per token prices unchanged.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source 2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source Google's guidance is explicit: for efficiency first workloads, stay on 3.7 Flash, which remains supported.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source

Then the schedule. The 0.75 and 3.75 dollar rates last through 31 December 2026; from 1 January 2027 both double.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source A task that costs 0.58 dollars today at high reasoning costs about 1.16 in January, and about 1.6 times what the same task cost on 3.7 Flash, figures this outlet derives from the reported numbers. The bargain is real and it is dated.

Who gains and who loses

Developers building coding and finance agents gain frontier tier results at a fraction of Opus 5's output price for the rest of 2026, and Android developers gain it inside their IDE.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source 3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source Google gains a fourth release in four months that keeps Flash in every price comparison. The Fairwind Program's defenders gain Flash Cyber, which the Chrome Security team says produces 2.6 times more correct patches than competing models and which posts a vulnerability discovery rate above 70 percent.3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source 4 DataCamp 2026-09-02 Terminal-Bench 2.1 at 90.8 vs 81.6 in this account; SWE-Bench Pro 61.6; SWE-Atlas 51.9; HLE flat at 45.4 vs 45.7; Flash Cyber vulnerability discovery above 70 percent, CWE-Bench 47.2; Fairwind gating; defensive bias. Open source

Anthropic loses the price comparison on the benchmarks where Flash ties Opus 5, and OpenAI faces a Flash at 59 on the index level with GPT-5.6 Sol.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source Latency sensitive applications lose: a 13 second first token is unusable for interactive chat.3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source Non English users absorb a slight safety regression the model card acknowledges.3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source And anyone who prices a product on today's rate loses in January.

The counter case

The claim that the effective price triples assumes buyers run at high reasoning. At low reasoning the model costs 0.24 dollars per task and finishes in 0.8 minutes, and if the intelligence gain survives at that setting the deal holds even after the doubling.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source The token inflation is also a choice Google exposed rather than hid, with three effort levels and an explicit recommendation.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source On capability, the OSWorld gap could be read as evidence that the coding results are narrow, but coding agents are where the money is and Google optimized accordingly. The benchmark figures are Google's, and one outlet's Terminal-bench number disagrees with another's, which is a reminder that vendor tables should be read as claims until reproduced.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source 4 DataCamp 2026-09-02 Terminal-Bench 2.1 at 90.8 vs 81.6 in this account; SWE-Bench Pro 61.6; SWE-Atlas 51.9; HLE flat at 45.4 vs 45.7; Flash Cyber vulnerability discovery above 70 percent, CWE-Bench 47.2; Fairwind gating; defensive bias. Open source

What to watch

  • The January price actually doubles. If Google holds the 1.50 and 7.50 schedule on 1 January 2027, the promotion was a promotion; an extension would signal Flash is losing share at full price.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source
  • A 3.9 or 4.0 Flash before year end. A fifth Flash release within four months would confirm the cadence and make the 3.8 price question moot before it bites.1 Artificial Analysis 2026-09-02 Intelligence Index 59, up 3, level with GPT-5.6 Sol sub maximum and Grok 4.6; 48,000 output tokens per task, up 30 percent; cost per task up about 40 percent to 0.58 dollars at high; 2.5 minutes per task; about 300 tokens per second; 0.75 and 3.75 dollars through year end then 1.50 and 7.50. Open source
  • Independent OSWorld results. Third party computer use scores near 59 percent would confirm the weakness is real; a materially higher number would suggest Google underclaimed.2 CellCog 2026-09-02 GA channels; pricing schedule to 1 January 2027; context 1,048,576 and 65,536 output; Google benchmark table versus 3.7 Flash and Opus 5; advice to stay on 3.7 for cost sensitive work; Gray Swan 5.5 from 9.2 percent; Flash Cyber via Fairwind, CyberGym 86.2. Open source
  • Flash Cyber patch results in the wild. Published counts of Chrome or open source vulnerabilities patched with Flash Cyber assistance by early 2027 would validate the 2.6 times claim.3 eesel AI 2026-09-02 HLE-Verified 54.9 vs 54.4, Harvey Legal Agent 10.0 vs 6.7; 17th of 196 on intelligence; third on speed; time to first token 13.30 seconds against 2.99 median; verbosity 120 million tokens against 71 million; built on 3.7 Flash; non English safety regression; Chrome Security 2.6 times more correct patches. Open source

Gemini 3.8 Flash is the same engine with the throttle open and a coupon attached. Buyers should price the engine, not the coupon.