DeepSeek moved its V4-Flash model out of preview on July 31, 2026, releasing the official API build, designated DeepSeek-V4-Flash-0731, in public beta. Artificial Analysis reported that the new build scores 50 on the Artificial Analysis Intelligence Index v4.1, a 10-point jump over the April preview, placing it second among all 162 models measured, and six points ahead of the larger DeepSeek V4-Pro on the same index. The result is a rare case of a smaller model overtaking its bigger sibling through post-training alone.
The release is notable not for a price change, pricing remains unchanged at $0.14 per million input tokens and $0.28 per million output tokens, with a 98% cache-hit discount, but for what the performance gains reveal about the limits of raw model size as a competitive differentiator. DeepSeek’s own changelog is explicit: the architecture and parameter count are identical to the April preview. Every improvement comes from re-post-training.
A Smaller Model Beating a Bigger One
The V4-Flash-0731 build carries the same 284-billion-parameter Mixture-of-Experts architecture as the April preview, with 13 billion parameters active at inference time and a one-million-token context window. What changed is its behavior on long-horizon agentic tasks. On Terminal Bench 2.1, the model scores 82.7, above the 76.1 that Kimi K3 posted at its launch two weeks ago, which had itself been leading every proprietary model tracked by independent evaluators. On DeepSWE, it scores 54.4; on Toolathlon, 70.3.
The agentic Elo on GDPval-AA v2, Artificial Analysis’s evaluation for real-world work tasks, rose from 1,189 for the April preview to 1,559 for the 0731 build. That places it second among open-weight models, behind Kimi K3’s 1,687 but ahead of GLM-5.2’s 1,510, a ranking that would have been unthinkable for a Flash-tier model just three months ago.
DeepSeek’s own framing is pointed: the 0731 build “far exceeds V4-Pro-Preview” on agent benchmarks, at roughly a third of the output price. The V4-Pro official release is still pending, but the Flash model’s performance advantage in agentic work means developers no longer have a clear reason to wait for it.
The Pricing Context
The output price of $0.28 per million tokens is unchanged from the April preview and is approximately one-third of V4-Pro’s $0.87 per million output tokens. The 98% cache-hit discount, $0.003 per million tokens on cached input, is among the most aggressive in the industry, significantly undercutting the 90% discount offered by most competitors.
Artificial Analysis notes one important caveat: the model is verbose, generating approximately 206 million output tokens across its evaluation suite against a median of 62 million for comparable models. That verbosity means real-world costs run meaningfully higher than the per-token sticker price implies, and developers running high-volume agentic workloads should budget on tokens generated rather than price per token.
The 0731 build also introduces native Responses API support and is specifically adapted for Codex, meaning developer workflows built on OpenAI’s Responses format work without an adapter, a practical integration advantage that lowers the switching cost for teams already running OpenAI-compatible tooling.
The Open-Weight Commitment
DeepSeek has confirmed that the full model weights will be released publicly in the coming weeks under an MIT license, consistent with the company’s pattern of open-sourcing its flagship models after the API release. This commitment to open weights is a defining feature of DeepSeek’s competitive strategy and one that has significant geopolitical implications: once the weights are live, any developer anywhere in the world can run a model that now scores second on the world’s most comprehensive AI intelligence index at zero licensing cost.
The V4-Flash-0731 release is best understood not as a pricing event but as a benchmark event. It demonstrates that post-training, the process of refining a model’s behavior through reinforcement learning and instruction tuning after initial pretraining, can deliver performance gains that rival or exceed those from scaling up model size. For a company that has consistently achieved more with less, it is a characteristically DeepSeek result. As we reported in our analysis of China’s three-punch open-source AI strategy, the pattern of Chinese labs releasing powerful open-weight models that outperform proprietary alternatives is now well established, and V4-Flash-0731 is the latest and most striking example yet.
The broader competitive implication is straightforward: if a 284-billion-parameter model can beat a 1.6-trillion-parameter model on the tasks that matter most to enterprise developers, the economics of the AI industry shift fundamentally. Scale still matters for some workloads, but the assumption that bigger always means better, and that premium pricing is therefore justified, is harder to sustain with each successive DeepSeek release. The V4-Flash-0731 result is a direct challenge to that assumption and a preview of how the open-weight era is likely to unfold: not through a single dramatic breakthrough, but through the steady, methodical accumulation of post-training gains that close the gap with proprietary models one benchmark at a time.
