{"slug": "inference-is-the-last-llm-moat", "title": "Inference Is the Last LLM Moat", "summary": "OpenAI and Anthropic's remaining competitive moat is subsidized inference, and losing it would cost them 50% or more of individual developer customers and small businesses, according to an analysis published on vroni.com. The author, who holds multiple OpenAI and Anthropic subscriptions, reports switching to DeepSeek v4.1 Flash via pi and OpenRouter once weekly usage limits are exhausted, citing OpenRouter provider unreliability, high cached-token prices, and timeouts. The piece argues that running inference at scale is the only real moat left for OpenAI and Anthropic, and that LLM inference will eventually become the new shared webhosting as other providers catch up.", "body_md": "### Built-In Tools vs Custom Tools in LLM Agents\n\nWhen building AI agents, a tool can be anything that the model can ask to use such as a search engine, a…\n\nOpenAI and Anthropic's moat for developers is subsidized inference, meaning reliable, fast, Western-hosted access to running models at an affordable monthly price. If that breaks, I think they will immediately lose 50%+ of individual developer customers and small businesses, large corporations will follow.\n\nTheir business looks solid right until it breaks. If OpenAI and Anthropic stop subsidizing heavy usage, people will flee overnight to DeepSeek and whatever good open model is cheapest that week.\n\nI have multiple OpenAI and Anthropic subscriptions and whenever I run out of weekly usage on all of them, I switch to DeepSeek, using it with pi and OpenRouter. The main annoyance here is that the providers on OpenRouter are a mixed bag, some of them I don't trust with my data, some of them have too high prices, especially for cached tokens, others are just not reliable, producing timeouts and other failures. While DeepSeek v4.1 Flash is clearly not on par with GPT-6-Astra or Fable 5.1, it is good enough as a daily driver. The main reason why I still prefer using OpenAI and Anthropic is because they are fast, reliable and cheap enough when using their subsidized subscriptions.\n\nRunning inference at scale is hard, and no one is as experienced with this as OpenAI and Anthropic. This is the only real moat they still have. But it's only a matter of time until we will see other providers catching up. At some point LLM inference will be the new shared webhosting.\n\nGive Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.\n\nTake a look at vroni.com", "url": "https://wpnews.pro/news/inference-is-the-last-llm-moat", "canonical_source": "https://www.vincentschmalbach.com/inference-is-the-last-llm-moat/", "published_at": "2026-09-16 03:48:16+00:00", "updated_at": "2026-09-16 04:09:00.411102+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-agents", "ai-tools", "ai-startups"], "entities": ["OpenAI", "Anthropic", "DeepSeek", "DeepSeek v4.1 Flash", "OpenRouter", "pi", "GPT-6-Astra", "Fable 5.1"], "alternates": {"html": "https://wpnews.pro/news/inference-is-the-last-llm-moat", "markdown": "https://wpnews.pro/news/inference-is-the-last-llm-moat.md", "text": "https://wpnews.pro/news/inference-is-the-last-llm-moat.txt", "jsonld": "https://wpnews.pro/news/inference-is-the-last-llm-moat.jsonld"}}