Member-only story
So DeepSeek just did something that genuinely made me stop scrolling.
On July 31 they shipped DeepSeek-V4-Flash-0731. It’s the same Flash model they released back in April. Same size. Same architecture. Same price $0.14 per million input tokens, which is roughly the cost of a nice coffee for a whole day of heavy use.
They changed exactly one thing: the training that happens after the model has read the internet.
And now it beats their own flagship — a model reported at 1.6 trillion parameters — on every single agent benchmark DeepSeek published. All nine of them.
The cheap one beat the expensive one. At the job everyone actually cares about in 2026: running agents.
I saw people talking about this all weekend. Some got it, some shrugged, and some jumped straight to “benchmarks are rigged.” So let’s clear things up because whichever camp you’re in, this release quietly breaks a rule most of us have been following for years without ever saying it out loud.
The Rule We All Believed #
Here’s the rule:
bigger model, better results.