DeepSeek put an experimental multimodal model, deepseek-v4-flash-vision-exp, on its API platform on 21 August 2026, describing it as matching the text only DeepSeek-V4-Flash on agents, reasoning and world knowledge while bringing multimodal agent performance close to Opus-4.8.1 DeepSeek API Docs 2026-08-21 DeepSeek's release note: deepseek-v4-flash-vision-exp live on the API platform on 21 August 2026; matches V4-Flash on text capabilities including agents, reasoning and world knowledge; multimodal agent performance close to Opus-4.8; one image capped at 384 tokens at V4-Flash pricing; Chat Completions, Messages and Responses support; base64, URL and free Files API image input; Harness 0.1.1 shipped same day with out of the box support. Open source The stake is not the vision capability itself, which is now table stakes, but the price at which it is being offered: roughly 87 cents to process a million words against roughly $50 from Anthropic, at benchmark scores that trail by low single digits on some tasks.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source We assess with moderate confidence that this release is a pricing action dressed as a capability release, and that its real effect will be measured in which agent workloads move off frontier priced APIs rather than in any benchmark table.

What the numbers actually say

DeepSeek published a comparison table rather than a claim, which makes the release unusually checkable. On ApexBench pass@1 the vision model scores 36.5 against 26.2 for the text only V4-Flash and 39.4 for Opus 4.8. On Agents' Last Exam it scores 27.3 against 25.7 for Opus 4.8, and on ZeroBench pass@5 it scores 35.0 against 34.0. On Chartography it scores 64.3 against 65.0, and on Terminal Bench 2.1 it reaches 83.9 against 82.7 for V4-Flash and 85.0 for Opus 4.8.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source The pattern is trading wins on narrow multimodal agent tasks and losing on the harder repository scale work: NL2Repo lands at 57.7 against 69.7.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source The Next Web counted the full table and found DeepSeek ahead on three of eleven benchmarks, with roughly twelve point deficits on NL2Repo and DSBench-Hard.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source

Three caveats sit on top of those figures and none of them are hidden. DeepSeek ran the evaluations in its own Harness Minimal Mode, with max tokens at the ceiling, top_p at 0.95 and temperature at 1.0, and the scores have not been independently reproduced.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source The text only V4-Flash column on multimodal benchmarks reflects a model that simply ignores the visual input, which inflates the apparent improvement from adding vision.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source And the comparison target, Opus 4.8, is no longer Anthropic's frontier model.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source Read together, the honest summary is narrower than the headline: an experimental model closes to within about three points of a superseded frontier model on a subset of multimodal agent tasks, on the vendor's own scoring rig.

The architecture is the tell

This is not a from scratch multimodal build. The model extends the existing text based V4-Flash with image understanding on top, and DeepSeek presents continuity with the base model as a feature rather than a limitation.2 The Decoder 2026-08-21 V4-Flash-Vision-Exp extends the text based V4-Flash with image processing rather than being a from scratch multimodal build; targets image description, screenshot text extraction and diagram analysis; up to 600 images per request; maximum edge 8,192 pixels falling to 4,096 with fifteen or more images; normalization to roughly 800 by 800; optional detail field downscaling to 512 by 512; free Files API to 64 MiB against 32 MiB for direct embedding. Open source The Next Web notes the awkward consequence: on text only benchmarks the vision variant beat its predecessor on six of seven tests, which contradicts the company's framing that it merely matches prior text performance.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source Either the vision training improved the text model, or the two runs are not strictly comparable. DeepSeek has not said which.

The engineering choices around the image path point the same direction, at cost rather than ceiling. Every image is capped at 384 tokens regardless of resolution, normalized to roughly 800 by 800 pixels, with an optional detail field that downscales to 512 by 512 to save tokens further.2 The Decoder 2026-08-21 V4-Flash-Vision-Exp extends the text based V4-Flash with image processing rather than being a from scratch multimodal build; targets image description, screenshot text extraction and diagram analysis; up to 600 images per request; maximum edge 8,192 pixels falling to 4,096 with fifteen or more images; normalization to roughly 800 by 800; optional detail field downscaling to 512 by 512; free Files API to 64 MiB against 32 MiB for direct embedding. Open source A request can carry up to 600 images, with the maximum edge length falling from 8,192 pixels to 4,096 once fifteen or more images are attached.2 The Decoder 2026-08-21 V4-Flash-Vision-Exp extends the text based V4-Flash with image processing rather than being a from scratch multimodal build; targets image description, screenshot text extraction and diagram analysis; up to 600 images per request; maximum edge 8,192 pixels falling to 4,096 with fifteen or more images; normalization to roughly 800 by 800; optional detail field downscaling to 512 by 512; free Files API to 64 MiB against 32 MiB for direct embedding. Open source Alongside it DeepSeek shipped a free Files API so a single uploaded image can be referenced across many requests instead of re transmitted, with a 64 MiB ceiling against 32 MiB for direct embedding.1 DeepSeek API Docs 2026-08-21 DeepSeek's release note: deepseek-v4-flash-vision-exp live on the API platform on 21 August 2026; matches V4-Flash on text capabilities including agents, reasoning and world knowledge; multimodal agent performance close to Opus-4.8; one image capped at 384 tokens at V4-Flash pricing; Chat Completions, Messages and Responses support; base64, URL and free Files API image input; Harness 0.1.1 shipped same day with out of the box support. Open source 2 The Decoder 2026-08-21 V4-Flash-Vision-Exp extends the text based V4-Flash with image processing rather than being a from scratch multimodal build; targets image description, screenshot text extraction and diagram analysis; up to 600 images per request; maximum edge 8,192 pixels falling to 4,096 with fifteen or more images; normalization to roughly 800 by 800; optional detail field downscaling to 512 by 512; free Files API to 64 MiB against 32 MiB for direct embedding. Open source These are the parameters of a system designed for long running screen reading agents, where the same interface is looked at hundreds of times, not for high fidelity document or medical imaging. We assess with moderate confidence that the target workload is computer use and browser automation.

Second order effects and the ledger

The pricing is the mechanism through which any of this matters. Images bill at V4-Flash rates, and V4-Flash off peak cache miss input runs $0.22 per million tokens with output at $0.66, doubling to peak rates during two windows on weekdays.5 DeepSeek API Docs, Models and Pricing 2026-08-21 deepseek-v4-flash listed with 1M token context and 384K maximum output as the cheapest model in the line; deepseek-v4-flash-vision-exp priced at the same rates; peak hours 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays with peak input at double off peak; off peak cache miss input $0.22 per million tokens for flash against $0.66 for v4-pro; flash output $0.66 per million off peak. Open source 3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source A vision agent that would otherwise burn its budget on image tokens is now capped at 384 tokens per frame on the cheapest model in the line.5 DeepSeek API Docs, Models and Pricing 2026-08-21 deepseek-v4-flash listed with 1M token context and 384K maximum output as the cheapest model in the line; deepseek-v4-flash-vision-exp priced at the same rates; peak hours 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays with peak input at double off peak; off peak cache miss input $0.22 per million tokens for flash against $0.66 for v4-pro; flash output $0.66 per million off peak. Open source

Who gains. Builders of high volume, low stakes visual agents gain the most directly: screenshot driven QA, web scraping with layout awareness, document triage where a wrong answer is caught downstream. At a roughly fifty to one cost ratio against Anthropic, a three point benchmark deficit is not a purchasing objection for those workloads.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source DeepSeek itself gains an evaluation narrative it can run on its own terms, since publishing a table against a competitor's previous generation sets the frame before anyone reproduces it. Chinese cloud and agent tooling vendors gain a domestic multimodal endpoint at commodity pricing, and Harness, DeepSeek's own agent framework, gains distribution by shipping same day support.1 DeepSeek API Docs 2026-08-21 DeepSeek's release note: deepseek-v4-flash-vision-exp live on the API platform on 21 August 2026; matches V4-Flash on text capabilities including agents, reasoning and world knowledge; multimodal agent performance close to Opus-4.8; one image capped at 384 tokens at V4-Flash pricing; Chat Completions, Messages and Responses support; base64, URL and free Files API image input; Harness 0.1.1 shipped same day with out of the box support. Open source

Who loses. Anthropic loses the low end of the vision agent market rather than the top, because the benchmarks where Opus 4.8 keeps a real lead are the repository scale and hard data science tasks that enterprises pay frontier prices for in the first place.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source Mid tier multimodal API vendors with neither a frontier score nor a commodity price lose the clearest, since this release removes the space between those two positions. And any buyer relying on published benchmark tables loses signal: a vendor scored table against a superseded competitor model, run on the vendor's own harness, is marketing with numbers in it until someone reproduces it.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source

The counter-case

The strongest argument against reading this as a pricing action is that the capability jump on multimodal agent tasks is real and large in its own frame: ApexBench moves from 26.2 to 36.5 within the same model family.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source If that transfers to work outside the benchmark suite, DeepSeek has closed a capability gap and the price is incidental. A second objection is that the word experimental in the model name is doing real work. This shipped as an exp endpoint, not a general availability model, and experimental endpoints get deprecated, re priced, or rate limited without notice, so no cost comparison against a production Anthropic model is stable. A third is that the fifty to one price ratio comes from a single outlet's calculation rather than a like for like published comparison, and cache hit rates, peak hour surcharges and retry behavior on a weaker model can compress a headline ratio substantially in production.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source 5 DeepSeek API Docs, Models and Pricing 2026-08-21 deepseek-v4-flash listed with 1M token context and 384K maximum output as the cheapest model in the line; deepseek-v4-flash-vision-exp priced at the same rates; peak hours 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays with peak input at double off peak; off peak cache miss input $0.22 per million tokens for flash against $0.66 for v4-pro; flash output $0.66 per million off peak. Open source For our reading to be wrong, the capability gain has to hold up under independent evaluation and the endpoint has to survive into general availability at these rates.

What to watch

  • Independent reproduction of the ApexBench and ZeroBench numbers. If third party evaluators reproduce scores within roughly two points of DeepSeek's published table within three months, the Harness Minimal Mode caveat can be retired.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source Silence or a large gap would mean the table was a marketing artifact.
  • Whether the exp suffix comes off. A general availability vision model at unchanged V4-Flash rates before the end of 2026 confirms the pricing thesis.5 DeepSeek API Docs, Models and Pricing 2026-08-21 deepseek-v4-flash listed with 1M token context and 384K maximum output as the cheapest model in the line; deepseek-v4-flash-vision-exp priced at the same rates; peak hours 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays with peak input at double off peak; off peak cache miss input $0.22 per million tokens for flash against $0.66 for v4-pro; flash output $0.66 per million off peak. Open source A quiet deprecation or a separate higher vision rate card would mean the 384 token cap was not economically sustainable.
  • A comparison against Opus 5 rather than Opus 4.8. DeepSeek's next release note either benchmarks the current frontier model or it does not.4 The Next Web 2026-08-21 DeepSeek won three of eleven published benchmarks against Opus 4.8 and trails by roughly twelve points on NL2Repo and DSBench-Hard; the comparison uses Opus 4.8 rather than the newer Opus 5; the text only V4-Flash ignores multimodal elements on certain tests; the vision model beat its predecessor on six of seven text only tests; processing one million words costs roughly 87 cents from DeepSeek against roughly $50 from Anthropic. Open source Continuing to score against a superseded generation is the clearest available signal that the real gap is wider than the published one.
  • Movement on the repository scale benchmarks. NL2Repo at 57.7 against 69.7 is the number that keeps enterprise coding agents on frontier priced APIs.3 OfficeChai 2026-08-21 Published benchmark table: Terminal Bench 2.1 83.9 vs 82.7 vs 85.0; NL2Repo 57.7 vs 69.7; ApexBench pass@1 36.5 vs 26.2 vs 39.4; Agents' Last Exam 27.3 vs 25.7; ZeroBench pass@5 35.0 vs 34.0; Chartography 64.3 vs 65.0. Evaluated in DeepSeek's Harness Minimal Mode at top_p 0.95 and temperature 1.0, not independently verified; text only V4-Flash scores reflect a model ignoring visual input; images billed at up to 384 tokens at V4-Flash rates with the Files API free. Open source Closing half that gap in the next model generation would be the point at which the price argument reaches the workloads that actually carry margin.
  • Whether the 384 token image cap becomes an industry pattern. If other providers ship hard per image token ceilings and free file reuse APIs within six months, DeepSeek will have set the cost structure for visual agents rather than merely competed on it.1 DeepSeek API Docs 2026-08-21 DeepSeek's release note: deepseek-v4-flash-vision-exp live on the API platform on 21 August 2026; matches V4-Flash on text capabilities including agents, reasoning and world knowledge; multimodal agent performance close to Opus-4.8; one image capped at 384 tokens at V4-Flash pricing; Chat Completions, Messages and Responses support; base64, URL and free Files API image input; Harness 0.1.1 shipped same day with out of the box support. Open source 2 The Decoder 2026-08-21 V4-Flash-Vision-Exp extends the text based V4-Flash with image processing rather than being a from scratch multimodal build; targets image description, screenshot text extraction and diagram analysis; up to 600 images per request; maximum edge 8,192 pixels falling to 4,096 with fifteen or more images; normalization to roughly 800 by 800; optional detail field downscaling to 512 by 512; free Files API to 64 MiB against 32 MiB for direct embedding. Open source

The interesting question this release leaves open is not whether a cheap Chinese model can approach a frontier one on a benchmark. It is what happens to the frontier price when the workloads that consume the most tokens, agents watching screens all day, stop needing the frontier at all.