DeepSeek’s $0.14 Model Just Beat Its Own Flagship at Agent Work DeepSeek released DeepSeek-V4-Flash-0731 on July 31, a $0.14 per million input tokens model that outperforms its own flagship, a 1.6 trillion parameter model, on all nine agent benchmarks published by DeepSeek. The update only changed post-training, challenging the assumption that larger models always yield better results. Member-only story DeepSeek’s $0.14 Model Just Beat Its Own Flagship at Agent Work So DeepSeek just did something that genuinely made me stop scrolling. On July 31 they shipped DeepSeek-V4-Flash-0731. It’s the same Flash model they released back in April. Same size. Same architecture. Same price $0.14 per million input tokens, which is roughly the cost of a nice coffee for a whole day of heavy use. They changed exactly one thing: the training that happens after the model has read the internet. And now it beats their own flagship — a model reported at 1.6 trillion parameters — on every single agent benchmark DeepSeek published. All nine of them. The cheap one beat the expensive one. At the job everyone actually cares about in 2026: running agents. I saw people talking about this all weekend. Some got it, some shrugged, and some jumped straight to “benchmarks are rigged.” So let’s clear things up because whichever camp you’re in, this release quietly breaks a rule most of us have been following for years without ever saying it out loud. The Rule We All Believed Here’s the rule: bigger model, better results.