{"slug": "why-we-wont-see-another-deekseek-moment-anytime-soon", "title": "Why we won’t see another DeekSeek moment anytime soon", "summary": "Moonshot AI's open-weight Kimi K3 model, released in mid-2025, ranks near the top of intelligence benchmarks but suffers from severe performance issues, including throughput dropping from 30 to 13 tokens per second and time to first token exceeding 20 seconds, forcing the company to pause new subscriptions. The model's high compute demands—requiring up to $4 million in GPUs per rack—contradict the market's earlier assumption, triggered by DeepSeek's January 2025 release, that open models would reduce the need for compute. Instead, cheaper and more capable frontier models have increased compute demand, with frontier pricing falling 4-5x to $15-$30 per million tokens while demand grew by three orders of magnitude and capability rose 32x, supporting Satya Nadella's Jevons paradox observation.", "body_md": "# Why we won’t see another DeekSeek moment anytime soon\n\nIn January 2025 the world saw the stock market collapsing with the ‘DeepSeek moment’. What happened: DeepSeek released an open weight model making a big leap in capabilities compared to predecessors. The result: NVIDIA lost about $600B in value on a single day. The logic from the market was: if great models can be made cheap and open, who needs that enormous amount of compute? Now, eighteen months later, recently released open models like GLM-5.2 and Kimi K3 are giving a similar shockwave to the AI market, but my prediction is that this time the stock market won’t collapse, more like the opposite.\n\nKimi K3 hit #8 on the OpenRouter leaderboard within a few days after release, doing ~155B tokens per day. And also [on the benchmarks](https://artificialanalysis.ai/models), the model performs exceptionally well, with only Claude Fable 5 and GPT-5.6 having a higher intelligence score on Artificial Analysis. We are seeing a similar leap in capabilities at Kilo.\n\nSo all of that is interesting, but aside from intelligence there’s a lot more to take into consideration. And these two charts tell an interesting story:\n\nTo understand the effectiveness of a model, you need to combine intelligence with cost and performance. And when you do that, the picture starts to look very different for Kimi K3: the model looks stellar on raw intelligence, but isn’t even to be found on the speed chart. That’s why, for [KiloBench](https://kilo.ai/kilobench), we’re looking at cost vs performance (in the Kilo harness) and popularity combined.\n\nThis week, Moonshot AI posted [this](https://x.com/Kimi_Moonshot/status/2078855608565207130):\n\nThe takeaway: even though the model is high on intelligence, performance metrics fell off a cliff. Throughput went from 30 tokens per second to 13. Time to first token increased to 20+ seconds. Moonshot paused new subscriptions to protect existing users and started splitting its plans to divide capacity between chat and coding.\n\nSo a frontier-class open model launched, and instead of relieving pressure on compute, it did the opposite. The DeepSeek moment made people believe that open models would make compute worthless, when in reality these high intelligence models like Kimi K3 show that it might actually be the only thing that matters.\n\n# Cheaper models ignite demand\n\nFrontier pricing has fallen from around $60 per million tokens three years ago to somewhere between $15 and $30 today, a 4-5x decline. In that same window, demand for frontier intelligence grew by at least three orders of magnitude, and measured capability (how long a task a model can complete on its own) went up roughly 32x by METR’s tracking.\n\nSo put simply, every time the frontier gets better and a little cheaper, the world finds a ton more to do with it. Cheaper tokens don’t reduce the bill, they expand the set of work worth running until you end up using more compute than before. Satya Nadella called this Jevons paradox the week DeepSeek hit, so that’s not the interesting part anymore.\n\n# The scarce thing was never the model\n\nBefore DeepSeek, everyone assumed the closed frontier model was the moat. After DeepSeek, the assumption flipped: open models commoditize the frontier, so the moat is gone. And both are wrong. Kimi K3 is open and near the top of the intelligence charts, but many people underestimate how much compute-backed capacity it takes to actually serve a model with this many parameters.\n\nAnd that compute capacity is expensive and slow to build. As Menlo Ventures partner [Deedy Das shared](https://www.linkedin.com/feed/update/urn:li:activity:7484641226525286400/), serving even a quantized 2.8-trillion-parameter model like K3 runs roughly $500K in GPUs at the low end, and closer to $4M for the rack you’d actually want. Neocloud providers are signing 3-5 year commitments that require serious upfront payments, and the market is taking them. Smaller buyers supposedly get sent away.\n\nOpen weight models are great, but you still need the hardware to host them competitively.\n\n# Capacity is scarce and volatile\n\nThe part that should worry anyone building on top of these models: it’s not just compute that is tight, capacity as a whole moves under your feet.\n\nSome model providers get throttled by their own success, but Claude Fable 5, the [best coding model on the market the week it launched](https://blog.kilo.ai/p/we-predicted-the-100kyr-per-dev-ai), was pulled for everyone by an export-control directive straight from the US government. And GitHub Copilot implemented usage-based billing that turned a fixed seat cost into a massive and unpredictable bill. What to take away from all of that: a model can get slower, pricier, or vanish, with no notice.\n\n# Why this points to routing\n\nIf capacity is the scarce and volatile thing, the winning move is making sure you’re vendor and provider agnostic.\n\nThe market is starting to realize this and look for an answer. Routing sends each task to whatever model fits, and increasingly to the provider that has capacity to serve it best. It has shown up across a wave of product launches lately, and the interest is picking up fast.\n\nEnterprises need the same thing with more at stake: the freedom to route across any model, governed and carried all the way to production. It only pays off, though, if the tooling has the freedom to route widely enough to matter. A layer locked to one provider solves nothing.\n\n# The crunch already arrived\n\nSo no, I don’t think we get another DeepSeek moment, at least not the version the market keeps bracing for. A cheap open model won’t make compute worthless, because the better and cheaper models get, the more compute the world wants. What’s coming instead is the capacity crunch.\n\nYou can’t control who wins the compute wars, so you better make sure that you’re not betting your whole workflow on a single model or provider.", "url": "https://wpnews.pro/news/why-we-wont-see-another-deekseek-moment-anytime-soon", "canonical_source": "https://blog.kilo.ai/p/no-second-deekseek-moment", "published_at": "2026-07-21 13:03:03+00:00", "updated_at": "2026-07-21 13:21:05.060451+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-chips", "ai-startups"], "entities": ["Moonshot AI", "Kimi K3", "DeepSeek", "NVIDIA", "OpenRouter", "Artificial Analysis", "KiloBench", "METR"], "alternates": {"html": "https://wpnews.pro/news/why-we-wont-see-another-deekseek-moment-anytime-soon", "markdown": "https://wpnews.pro/news/why-we-wont-see-another-deekseek-moment-anytime-soon.md", "text": "https://wpnews.pro/news/why-we-wont-see-another-deekseek-moment-anytime-soon.txt", "jsonld": "https://wpnews.pro/news/why-we-wont-see-another-deekseek-moment-anytime-soon.jsonld"}}