On July 31, 2026, DeepSeek published DeepSeek-V4-Flash-0731 to its official Hugging Face organization and updated its API documentation to route the deepseek-v4-flash model id to the new build. Existing API calls pick up the new version with no code change. The weights and repository are released under the MIT License.
The build is a re-post-training of the earlier V4-Flash preview rather than a new architecture, and the benchmark table DeepSeek published shows how much that alone moved. Against its own V4-Flash preview, Terminal Bench 2.1 rises from 61.8 to 82.7, NL2Repo from 39.4 to 54.2, DeepSWE from 7.3 to 54.4, Cybergym from 38.7 to 76.7, Toolathlon-Verified from 49.7 to 70.3, DSBench-FullStack from 37.0 to 68.7, and DSBench-Hard from 25.8 to 59.6. On all nine published benchmarks the smaller Flash build now scores above DeepSeek’s own larger V4-Pro preview. DeepSeek also compares it directly against GLM-5.2 and Claude Opus 4.8, where it leads GLM-5.2 on every listed benchmark and trails Opus 4.8 on all of them.
The company’s own framing of the gains is post-training focused on coding, agents, reasoning, and tool use, plus speculative decoding support and native compatibility with the Responses API format.
Two things matter here for buyers. First, DeepSeek is demonstrating that a large share of current agentic capability comes from post-training, not from a bigger model: same architecture, dramatically different scores. Second, an MIT-licensed model that beats the vendor’s own larger preview compresses the gap between what you can rent and what you can self-host, and it does so at a price point that puts continued pressure on Western API list prices.