DeepSeek V4-Flash 0731: The Moment the AI Landscape Shifted

Key Takeaways
Post‑training on a smaller MoE model can match or exceed larger frontier models on agent benchmarks.
$0.14 input / $0.28 output per million token pricing pressures closed labs.
Open‑source release accelerates community‑driven performance improvements and cost competition.
AI race shifts from massive pre‑training to efficient post‑training and pricing.
The morning the ground shifted again
I was scrolling X the way I always do — half-sleepy, half-expecting another quiet day in the AI timeline — when DeepSeek’s official account dropped the post.
One image. A clean table. Blue headers. Numbers that didn’t look real at first glance.
DeepSeek-V4-Flash-0731. Official API live in public beta. Same architecture. Same size. Just re-post-trained.
I stared at the table for a full minute.
That was the moment I felt it in my chest.
This is what we started calling a “DeepSeek moment.”
Not hype. Not a carefully staged demo. Just a quiet drop that forces every other lab to recalculate overnight.
I quote-tweeted it almost immediately:
This is how to show everyone why people started calling model releases a “DeepSeek moment.”
Just compare these values from the preview and you will be surprised with the massive gains in intelligence. Not too far behind Opus 4.8.
With this price, it’s a steal. I knew OpenAI’s discount will be the only possible strategy with which they could compete with DeepSeek V4 GA release.
I meant every word.
The numbers that rewrote the conversation
Here’s the exact image they posted — the one that made the timeline lose its mind:
DeepSeek’s official benchmark table
| Benchmark | V4-Flash-0731 | Preview | V4-Pro-Preview | GLM-5.2 | Opus-4.8 |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 |
| NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 |
| Cybergym | 76.7 | 38.7 | 52.7 | – | 83.1 |
| DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 |
| Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 |
| Agents’ Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 |
| AutomationBench (Public) | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 |
| DSBench-FullStack | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 |
| DSBench-Hard | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 |
Look at the jumps again, slower this time.
Terminal Bench 2.1 went from 61.8 → 82.7.
DeepSWE from a pathetic 7.3 → 54.4.
Cybergym nearly doubled.
Toolathlon-Verified, DSBench-FullStack, DSBench-Hard — all of them leapt.
These aren’t cherry-picked academic benchmarks. These are agent tasks. Terminal control. Software engineering. Cybersecurity environments. Full-stack development. The kind of work that used to require the most expensive frontier models.
And DeepSeek did it with pure post-training on a 284B total / 13B active MoE that was already months old. No new base model. No massive new pre-training run. Just better post-training. The smaller, cheaper model now outperforms its own bigger sibling on the metrics that matter for real agentic work. Artificial Analysis even put it at 50 on their Intelligence Index — a clean +10 point jump.
Then the weights dropped on Hugging Face a few hours later:
deepseek-ai/DeepSeek-V4-Flash-0731
MIT license. Speculative decoding (DSpark) already attached. The open-source inference community did what it always does — it moved. I said it out loud on X: give it two hours and you’ll see multiple providers on OpenRouter offering faster speeds than the official endpoint. That community is hyper-reactive, and it’s one of the healthiest forces in this entire industry.
The price that changes everything
$0.14 input / $0.28 output per million tokens.
Cache hit: $0.0028.
Concurrency limit: 2,500.
It natively supports the Responses API format and is fully adapted for Codex.
That is not “competitive pricing.” That is structural violence against every closed lab still charging five to thirty dollars for similar or worse agent performance.
I wrote it plainly earlier in the day:
With this price, it’s a steal. I knew OpenAI’s discount will be the only possible strategy with which they could compete…
OpenAI had just cut prices dramatically the day before. It felt like the only remaining lever they had left. DeepSeek answered the next morning by making even that lever look expensive.
Anthropic, meanwhile, continues to treat their Haiku-class models like the affordable audience was never real. That approach is starting to look like a strategic error of historical proportions. DeepSeek’s pricing at least keeps the pressure real.
What this actually means
I kept coming back to the same thought all day:
The game has completely switched.
The world will never be the same after August.
Mark my words.
Open models and truly cost-effective models are not a side story anymore. They are becoming the main story. The closed frontier labs in the United States are doomed if they keep relying on the current playbook — giant closed models, high prices, slow iteration, and the assumption that capability alone will keep customers locked in.
DeepSeek just showed that aggressive post-training on a smaller, efficient architecture can deliver agent performance that was supposed to belong only to the most expensive systems. And they did it while keeping the model open enough that the entire community can run with it.
Yes, it’s still text-only. Yes, the official V4-Pro is still coming. Those are real limitations. I even said as much in a quieter reply later in the day. But they don’t erase what just happened.
A smaller model. Pure post-training. Massive agent gains. Absurd pricing. Weights released the same day.
That is not incremental progress.
That is a different philosophy winning in public.
I don’t know exactly how the next few months will play out. But I know the old assumptions are dying faster than most people are willing to admit.
Today felt like one of those days you look back on later and say: that was when the ground finally moved.
Another DeepSeek moment.
By DeepSeek themselves.
Finished reading?
Connect with Gargeya Sharma on digital strategy and autonomous pipelines.