Skip to content

DeepSeek V4 Flash 0731 Review — A Budget Model That Beats Its Own Flagship

On July 31, 2026, DeepSeek pushed a public beta of DeepSeek V4 Flash under the build tag V4-Flash-0731, and the interesting thing about it is what did not change. Same architecture. Same parameter count. Same 1M-token context. Same $0.14 / $0.55 pricing as the April build. The only thing DeepSeek touched was post-training — and by its own agent benchmarks, that pass is enough to push the budget Flash model past the larger, pricier V4 Pro. This review covers what the re-training actually moved, which of the numbers are measured versus interpolated, and where V4 Flash 0731 fits.

TL;DR verdict

DeepSeek V4 Flash 0731
TypeOpen-weight budget model (agentic + coding focus)
Context window1,000,000 tokens
ModalitiesText
LicenseOpen weights
Pricing$0.14 in / $0.55 out per 1M — unchanged from the April V4 Flash
Headline numbersTerminal-Bench 2.1 82.7 · AA Intelligence Index 50 (+10 over prior Flash)
AvailabilityDeepSeek API, OpenRouter, Together, Fireworks; weights downloadable
Best forTool-heavy agents and coding pipelines on a tight budget
CaveatAgentic/coding gains are the published story — broad-reasoning cells are estimates

If you skip the rest: V4 Flash 0731 is one of the best intelligence-per-dollar picks for agentic work this month. DeepSeek re-trained a model it had already shipped and, without changing the architecture or the price, lifted it from a mid-tier flash model to one that — on the agent benchmarks it chose to publish — beats its own flagship. The honest asterisk is that this is a vendor’s sweep on vendor-selected agent evals, and the broad-reasoning placements are still interpolation.

What changed: post-training, not scale

The headline is a claim about method. V4-Flash-0731 is the same model as the April 2026 V4 Flash — identical architecture, identical parameter budget — with a new post-training run layered on top. That framing matters, because it separates two things labs usually ship together: a bigger or re-architected base model, and better alignment/RL on top of it. Here only the second moved, and the deltas DeepSeek reported are large:

  • DeepSWE (agentic software-engineering eval): 7.3 → 54.4, roughly a 7× jump in a single post-training cycle.
  • DSBench-FullStack (internal full-stack development test): 37.0 → 68.7.
  • Terminal-Bench 2.1: 82.7, up from V4-Pro-Preview’s 72.1 — a 14-point win for the cheaper model on agentic terminal work.

DeepSeek’s summary claim is that V4 Flash 0731 beats V4-Pro-Preview on all nine agent benchmarks it reported. Whether or not every one of those nine survives independent replication, the pattern is coherent: a small model that was leaving a lot of capability on the table, unlocked by RL tuned specifically for tool use and multi-step coding rather than one-shot exams.

What the numbers say

Here is where the case is strongest — and where it needs an asterisk. On the independent Artificial Analysis Intelligence Index, V4 Flash 0731 scores 50, a 10-point jump over the prior V4 Flash and about six points ahead of DeepSeek V4 Pro, landing one point behind GPT-5.6 Luna on that composite. That is a genuinely unusual result: a vendor’s budget tier out-scoring its own flagship on an independent blended index.

BenchmarkV4 Flash 0731Provenance
Terminal-Bench 2.182.7Vendor-reported (up from V4 Pro’s 72.1)
AA Intelligence Index50Independent aggregate (Artificial Analysis)
GPQA Diamond~91Aggregator-reported — recorded as an estimate
HLE~37Aggregator-reported — recorded as an estimate

The provenance column is the part to read carefully. The Terminal-Bench 2.1 figure is a first-party DeepSeek number. The Intelligence Index is an independent composite. But the GPQA Diamond and HLE placements come from third-party aggregators, not from a DeepSeek model card — so on our leaderboard they are recorded as estimates, greyed out and excluded from every ranking, anchored to DeepSeek V4 Pro’s vendor GPQA of 90.1. They are directionally useful; they are not measurements. What DeepSeek did not publish is the full standard suite (MMLU-Pro, SWE-bench Verified, MATH-500, MMMU-Pro), which is why the row still shows inherited legacy cells for those columns.

What it costs

Nothing. That is the point. Pricing is unchanged from the April V4 Flash — about $0.14 per 1M input and $0.55 per 1M output on DeepSeek’s API, with the usual hosted open-weight rates on OpenRouter, Together, and Fireworks. Artificial Analysis notes the 0731 build shares “identical architecture and pricing” with its predecessor, so the entire capability gain arrives at zero marginal cost to anyone already budgeting for V4 Flash.

Put that next to the closed frontier — Claude Opus 5 at $5/$25, GPT-5.6 Sol at $5/$30 — and the intelligence-per-dollar argument is stark: on the agentic axis DeepSeek chose to measure, a model two orders of magnitude cheaper on input is trading blows near the top. The cost calculator on the leaderboard lets you plug your own token mix to see where that crossover actually lands for your workload; for high-volume, tool-heavy pipelines it lands early.

How it compares

V4 Flash 0731’s real competition is the open-weight agentic tier, not the closed reasoning frontier. Against its own V4 Pro sibling it is the cheaper model that now wins the agent benchmarks — an odd but welcome inversion that makes Pro harder to justify for pure tool-use workloads. Against the open coding flagships — GLM-5.2 and Kimi K2.7-Code — V4 Flash 0731 is dramatically cheaper and lighter to serve, though those larger models still lead on the hardest coding evals; the trade is capability-per-parameter versus capability-per-dollar. Against the closed budget tier — Gemini 3.6 Flash, GPT-5.6 Luna — it undercuts on price and matches on the published agentic numbers while giving up multimodality, since V4 Flash is text-only.

The caution worth stating plainly: this is a vendor sweep on vendor-selected agent benchmarks, released as a beta. Independent evaluators have flagged elevated benchmark-gaming across the newest releases all year, and a model re-trained specifically to top agent evals is exactly the case where your own repository is the only benchmark that counts. The Intelligence Index result is independent and reassuring; the “beats Pro on all nine” claim is not yet.

Who should care

  • Teams running tool-heavy agents on a budget: This is the headline use case. A Terminal-Bench 2.1 of 82.7 at $0.14/$0.55 makes V4 Flash 0731 a serious default for function-calling and command-line agents. See multi-agent pipelines for where a cheap, strong tool-user slots into a larger workflow.
  • High-volume coding pipelines: If you are paying closed-frontier rates on a large PR-review or codegen pipeline, A/B this against your current coder on real pull requests before the next billing cycle — the price gap compounds fast.
  • Anyone who already deployed V4 Flash: The upgrade is a checkpoint swap at the same price. Re-run your eval set against 0731; the agentic gains may be free money.
  • Anyone ranking on broad reasoning: Wait for third-party standard-suite numbers before placing V4 Flash 0731 above frontier models on general knowledge. Its published strengths are agentic; its reasoning placement is an estimate. The LLM Benchmark Comparison 2026 covers how to weigh that trade.

FAQ

What is DeepSeek V4 Flash 0731? A re-trained checkpoint of DeepSeek’s open-weight V4 Flash, released as a public beta on July 31, 2026. Same architecture, parameter count, and 1M context as the April build; only the post-training changed. Priced at $0.14 in / $0.55 out per 1M.

Is it better than DeepSeek V4 Pro? On the agent benchmarks DeepSeek published, yes — Terminal-Bench 2.1 82.7 versus Pro’s 72.1, and a claimed sweep of nine agent evals. On the independent Intelligence Index it lands about six points ahead of Pro. Pro still leads on broad reasoning, and the two share a context window.

How much does it cost? Unchanged from April: about $0.14 per 1M input and $0.55 per 1M output on DeepSeek’s API, with hosted open-weight rates on the usual routers.

Does it have published benchmark scores? Agentic and coding ones — Terminal-Bench 2.1 82.7, DeepSWE 54.4 (up from 7.3), DSBench-FullStack 68.7 (up from 37.0). Its GPQA (~91) and HLE (~37) figures come from aggregators and are treated as estimates, not vendor measurements.

Is it open weights? Yes. V4 Flash is downloadable and self-hostable, and also available through the DeepSeek API, OpenRouter, Together, and Fireworks.

Continue reading

  • AI Models Leaderboard — DeepSeek V4 Flash 0731 versus 70+ models on benchmarks, pricing, and context, with a cost calculator.
  • GLM-5.2 Review — the open-weight coding flagship in the tier above, and how its published SWE-bench Pro numbers compare.
  • Kimi K2.7-Code Review — Moonshot’s open-weight coder, the other value pick in the same conversation.
  • Claude Opus 5 Review — the closed frontier leader V4 Flash 0731’s intelligence-per-dollar pitch is measured against.
  • LLM Benchmark Comparison 2026 — how to read agentic benchmarks and vendor sweeps without getting fooled.
  • All Reviews — index of every review on the site.