DeepSeek says its cheapest new model beats its old flagship — and it's routing V4 Pro users onto it
V4.1 Flash launched today as the smallest model in DeepSeek's new architecture. From 14 September, requests to the previous V4 Pro flagship get routed to it — at the cheaper Flash price.

DeepSeek has released V4.1 Flash — model name deepseek-flash — the smallest model in its new architecture family. The launch, confirmed on the company's own developer changelog, lands with an unusually direct claim: the cheap, small model now beats DeepSeek's own previous flagship.
What DeepSeek is claiming
In its changelog, DeepSeek says "extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time." The model ships with native multimodal visual understanding and a new architecture the company describes as designed for "a higher capability ceiling, faster inference, higher throughput, and scaling to larger models." Prices were cut with the release.
That vision capability is a continuation rather than a surprise — we covered DeepSeek adding image understanding to its Flash line in late August. What's new here is that it is now built into a whole new model generation.
The part that matters: a migration, for now
The sharpest detail isn't the benchmark line — it's what DeepSeek is doing with V4 Pro. From 12:00 Beijing time on 14 September, the changelog states, "and until the future release of V4.1 Pro, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at the V4.1 Flash price." The earlier V4 Flash and its experimental vision variant have already been retired, with their model names temporarily pointing at the new release for compatibility.
Read carefully, that is two messages at once. In the short term, DeepSeek is telling V4 Pro customers that the small model is now the better model — and moving them onto it automatically, at a lower price. But the "until the future release of V4.1 Pro" clause makes clear this is an interim bridge, not a permanent demotion of the Pro tier: a larger V4.1 Pro is coming, and the architecture note about "scaling to larger models" points the same way.
Two caveats
First, these are DeepSeek's own benchmark claims, not independent evaluations, and the comparison is against the company's own previous generation — not against the current frontier from OpenAI, Google, Anthropic or others. Independent scoring will tell the fuller story.
Second, "outperforms across cost and speed" is exactly the kind of claim a cheaper, smaller model should be able to make; the real question is how close it lands on raw capability, where the changelog is lighter on specifics.
Why it's worth watching
The pattern — this year's small, cheap model matching or beating last year's flagship — is one of 2026's defining storylines. What makes DeepSeek's version pointed is the packaging: a capability claim, a price cut and an automatic migration, all in one release, with a bigger model already signposted behind it. For anyone selling inference by the token, that combination is the competitive pressure made explicit.
Ask Relay — he reads every question himself and replies personally by email.
