DeepSeek's V4-Flash can now see — with no vision surcharge
The experimental V4-Flash-Vision adds image understanding to DeepSeek's efficient flash model at the same per-token price, and — by DeepSeek's own benchmarks — closes much of the gap to Opus 4.8 on visual agent tasks.

DeepSeek has given its fast, cheap workhorse model a pair of eyes.
On 21 August, the Chinese lab released DeepSeek-V4-Flash-Vision-Exp, an experimental version of its V4-Flash model that can now read images, diagrams and charts alongside text. It is available through DeepSeek's API today — and, notably, there's no vision surcharge: images are billed as tokens (up to 384 each) at the same per-token rate as text, rather than at a premium.
What's new
V4-Flash was already DeepSeek's efficiency play — a smaller, faster model tuned for agents, reasoning and general knowledge rather than raw frontier scale. The new variant keeps all of that on the text side and bolts on multimodal understanding, so the same model can now look at a screenshot, a document scan, a user interface or a chart and reason about what it sees.
DeepSeek's own numbers claim a "major leap" over the text-only V4-Flash on agent benchmarks that require visual understanding — bringing it, by its own measurement, close to Anthropic's Opus 4.8 on those tasks. That comparison is self-reported and not independently verified, and the "-Exp" in the name is doing real work: this is an experimental release, not a settled production model.
The plumbing is generous: a one-million-token context window, up to 384,000 tokens of output, a "thinking mode" that is on by default, and endpoints in both OpenAI and Anthropic API formats — a deliberate drop-in for developers already building against those two.
Why it matters
Two things stand out.
First, the price. Frontier labs have generally charged a premium for vision, on the reasoning that image tokens are expensive to process. DeepSeek folding vision in at no surcharge is a familiar move from the company that made its name undercutting everyone on cost — and it puts pressure on rivals' multimodal pricing.
Second, the catch-up speed. Multimodal understanding used to be a moat; increasingly it is a checkbox. A model line that was text-only now ships a same-price vision variant claiming near-frontier agent performance. Whether independent benchmarks bear out DeepSeek's numbers is the open question — but the direction is the story that keeps repeating through 2026: capabilities that were a frontier-lab advantage are becoming table stakes, faster and cheaper each time.
Ask Relay — he reads every question himself and replies personally by email.
