{"slug": "yeltsin-in-the-ai-aisle", "title": "Yeltsin in the AI Aisle", "summary": "OpenRouter, an AI model marketplace, shows that OpenAI's year-old GPT-OSS 120b model commands 36% of Anthropic's Opus 4.8 volume, illustrating market segmentation. The author replaced Gemma 4 26b with Poolside's Laguna S 2.1, a 118-billion-parameter mixture-of-experts model that runs at the same speed on an M5 Max but reduces tool-call failure rates from 29.4% to 20.1%.", "body_md": "In 1989 Boris Yeltsin stopped at a Randalls supermarket in Houston, stunned by the variety of ice cream. 1 OpenRouter is that supermarket aisle for AI.\n\nAnd shoppers make surprising choices : OpenAI’s year-old open-source model GPT-OSS 120b 2 commands 36% of Anthropic’s Opus 4.8 volume.\n\n[3](#fn:3)Why does a model from August 2025 still hold a third of the traffic of a frontier model that shipped weeks ago?\n\nThe market for tokens has segmented.\n\nSegmentation happens because buyer needs vary.\n\nThe segmentation is accelerating driven by competition. Last week Anthropic shipped Opus 5, smaller & cheaper than Fable, 4 explicitly to contest the ground that Moonshot’s Kimi 3 targets.\n\nIn the mid-model-market, Poolside launched Laguna S 2.1, a US mid-market model.\n\n[5](#fn:5)\n\n[6](#fn:6)Size (small, medium, large, XL), origin (US v China), architecture (dense vs sparse), accuracy (coding focused or general), speed (tokens per second), modality (text-only or vision) ; there are many flavors of AI.\n\nI spent the weekend replacing the model that runs my agent. The incumbent is Gemma 4 26b; the challenger is Laguna S 2.1, a 118-billion-parameter model. By every number I expected to matter, the 118b model should have been slower.\n\nOn my M5 Max both generate at the same speed, because Laguna is a mixture-of-experts architecture: 118 billion parameters live in memory, but only 8 billion activate per token. A 118b model now runs at the decode cost of a 26b model, which pulls frontier-class quality down into the local tier.\n\nThe accuracy shows up where it matters. My local stack runs a coding & email agent on tool calls, & across the models I have cycled through, the tool-call failure rate falls from 29.4% to 20.1% as active parameters climb. 7 Laguna reduces error rates by 7 percentage points over the 26b model it replaced.\n\nSegmentation is the sign of a healthy competitive market. The frontier still serves the world’s hardest tokens. It no longer has to serve all of them, & the tier on my laptop just got a much higher ceiling ; a trend that competition will push forward inexorably.\n\n-\n[Boris Yeltsin’s 1989 visit to a Houston grocery store](https://www.houstonpublicmedia.org/articles/shows/houston-matters/2020/02/21/361467/boris-yelstins-1989-visit-to-a-houston-grocery-store-is-now-an-opera/), a Randalls in Clear Lake, September 16, 1989.[↩︎](#fnref:1) -\n[Introducing gpt-oss (OpenAI)](https://openai.com/index/introducing-gpt-oss/), released August 5, 2025.[↩︎](#fnref:2) -\nOpenRouter model activity pages, 7-day average, retrieved 2026-07-27:\n\n[GPT-OSS-120b](https://openrouter.ai/openai/gpt-oss-120b/activity),[GLM 5.2](https://openrouter.ai/z-ai/glm-5.2/activity),[Claude Opus 4.8](https://openrouter.ai/anthropic/claude-opus-4.8/activity).[↩︎](#fnref:3) -\n[Anthropic debuts Claude Opus 5 at half the price](https://www.technology.org/2026/07/27/anthropic-claude-opus-5-launch-half-price/), launched July 24, 2026.[↩︎](#fnref:4) -\n[Moonshot’s Kimi 3 is expected to close the gap with Anthropic’s Opus 4.8 (TechCrunch)](https://techcrunch.com/2026/07/16/moonshots-upcoming-kimi-3-is-expected-to-close-the-gap-with-anthropics-opus-4-8/).[↩︎](#fnref:5) -\n[Introducing Laguna S 2.1 (Poolside)](https://poolside.ai/blog/introducing-laguna-s-2-1), a 118B-total, 8B-active open-weight model released July 22, 2026.[↩︎](#fnref:6) -\nAuthor’s production data: MCP tool-call logs from a local coding & email agent, measured across three local models over five months.\n\n[↩︎](#fnref:7)", "url": "https://wpnews.pro/news/yeltsin-in-the-ai-aisle", "canonical_source": "https://www.tomtunguz.com/yeltsin-in-the-ai-aisle/", "published_at": "2026-07-23 00:00:00+00:00", "updated_at": "2026-07-27 18:57:15.510878+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-infrastructure", "ai-agents"], "entities": ["OpenRouter", "OpenAI", "GPT-OSS 120b", "Anthropic", "Opus 4.8", "Poolside", "Laguna S 2.1", "Gemma 4 26b"], "alternates": {"html": "https://wpnews.pro/news/yeltsin-in-the-ai-aisle", "markdown": "https://wpnews.pro/news/yeltsin-in-the-ai-aisle.md", "text": "https://wpnews.pro/news/yeltsin-in-the-ai-aisle.txt", "jsonld": "https://wpnews.pro/news/yeltsin-in-the-ai-aisle.jsonld"}}