Is the future of AI locked-in? Frontier AI labs Anthropic and OpenAI are pursuing a lock-in strategy to retain users and recoup billions in training costs, evolving from model-level to organizational workflow-level integration, with the latest stage marked by Anthropic's Claude Tag for Slack. This lock-in, which now includes agents running on lab infrastructure, raises concerns for enterprises seeking to control costs, privacy, and IP. When you’re a frontier lab, i.e. Anthropic and OpenAI, and you’re spending billions of dollars to train AI models, you need to bring in as many users as you can and, most importantly, keep them; keep them using your models and keep increasing how much they use your models. In the first years of AI-as-a-Service, 2022-2024, bringing in users was the bigger challenge. Competition was a two horse race between OpenAI and Anthropic until January 2025 when the Deepseek moment occurred and Chinese models appeared not only as viable competitors from Chinese labs with deep talent pools, but the models were “open weight” – anyone could run them for themselves or serve them to other businesses. The frontier labs need to recoup the training costs of their models by serving them to as many people as possible. Serving models, inference, has very healthy margins https://www.forbes.com/sites/petercohan/2026/07/28/as-token-costs-plunge-enterprise-ai-providers-face-a-new-margin-squeeze/ possibly as high as 70% . The more tokens they can get users to burn the faster they close the gap between their immense compute build-out expenditures and their revenue generation. This has made lock-in the labs’ north star strategy. That lock-in has evolved over a few major stages as models have improved, and time and talent acquisition has allowed the supporting infrastructure to be put into place. They started with model lock-in. This was the early days where the interface was conversational and copy-and-paste was your IO: ChatGPT or Claude in the browser. In the background, businesses were calling APIs directly and making the first steps at integrating AI into their workflows. At the beginning of 2024 OpenAI began experimenting with memory for ChatGPT. In June Anthropic launched Artifacts https://madewithclaude.com/ and Projects. Now users could build with Claude while their files and what they built lived inside the interface. Lock-in advanced to the user and their work. In late 2024 Anthropic announced the Model Context Protocol https://modelcontextprotocol.io/ , which was rapidly adopted by industry as it let users integrate models with the services like Slack, Gmail and Google Drive that they were already using. Around that time OpenAI also released their Projects https://help.openai.com/en/articles/10169521-projects-in-chatgpt , which let your organise chats, reference files and instructions. This expanded lock-in to the ecosystem and data level, but was still limited to the individual. Work was still manually triggered. OpenAI launched Codex https://openai.com/index/introducing-codex/ , their coding agent, but running in their own cloud. It would take a few months before they followed Anthropic’s Claude Code into the terminal. Lock-in really started to take off in November 2025 when a step change in agentic ability https://davidw.tech/2026/03/21/no-longer-an-intern.html occurred with the release of Claude Opus 4.5. Long running workflows became feasible. What followed was the frontier labs creating the infrastructure to let users run those workflows in the labs’ clouds. The final stage is happening now, and it began with Claude Tag for Slack https://www.anthropic.com/news/introducing-claude-tag . Lock-in has reached the organisational and workflow level. Frontier lab agents will help your team by ingesting your data, linking your services with MCP, and memorising your instructions. You get more done, but you slowly grow more and more implicitly and explicitly reliant on frontier lab models running on frontier lab infrastructure. Escaping the lock-in via software On the other side of lock-in is enterprise working to control their costs, maintain privacy even outside of verticals where privacy is regulated , and protect their IP. Data is the new gold and no-one wants to hand it for free to labs that can train agents to exploit it. AI costs created two constrasting headlines recently. On the high end was Uber announcing they had burned through their annual token budget in 4 months https://www.forbes.com/sites/janakirammsv/2026/05/17/uber-burns-its-2026-ai-budget-in-four-months-on-claude-code/ . On the low end was Deepseek announcing in April this year their Deepseek v4 Pro https://www.deepseek.com/en/news/v4-preview/ model, which was within reach of the frontier labs top models, but was approximately 1/20th the price https://www.solvimon.com/pricing-guides/anthropic-vs-deepseek , but even cheaper in practice due to caching. And this model, and it’s smaller faster sibling, Deepseek V4 Flash, had weights available for anyone with the capable hardware these models were not small . Businesses have been building on top of open weight models since Facebook released their popular Llama 3 models in mid-2024. These were workhorses that weren’t replaced until the Qwen 3.5 line was released by Alibaba https://qwen.ai/blog?id=qwen3.5 in early 2026. 80% of US startups are now building on top of Chinese models https://www.newline.co/@Dipen/why-80percent-of-us-ai-startups-switched-to-chinese-models--9b216c28 . And AT&T announced they achieved savings of 80% to 90% https://about.att.com/blogs/2026/the-tokenomics-equation.html? by switching from the proprietary models of the frontier labs to open models running on their own hardware or in the cloud and by mixing in the intelligent routing of requests to the cheapest capable model. OpenRouter https://openrouter.ai/ , the AI world’s most popular middleman, which was recently bought by Stripe for $7B https://stripe.com/au/newsroom/news/stripe-agrees-to-acquire-openrouter , provides intelligent routing https://openrouter.ai/docs/guides/routing/routers/auto-router across its giant range of models from all the providers and neoclouds. And since most local AI harnesses including Codex and Claude Code plus all of the open source harnesses allow you to set the address you send requests to, every business can have this same feature with proper limits on providers to protect privacy and IP . Escaping the lock-in hardware Part of AT&T’s strategy is running models on their own hardware. Only in the last 3 months have the numbers changed to make this an option for businesses and individuals below the enterprise scale. What has changed the numbers is advances in inference and new models. Advances in inference have come from the ability of current frontier models Fable, GPT-5.6 and Kimi K3 to generate and optimise GPU code. Much like random Anthropic employees advancing the frontiers of math https://www.wsj.com/tech/ai/ai-math-riemann-hypothesis-anthropic-openai-22f98a87 , a random person can ask these models to make running a small LLM run faster on their hardware. And many have. These efforts combined with cutting edge research in speculative decoding https://en.wikipedia.org/wiki/Speculative decoding , and frontier models’ ability to turn research papers into working code, people are reporting running models locally at speeds competitive with inference providers or at the very least a usable speed with previously frontier lab only performance. The Qwen3.8 27B https://huggingface.co/Qwen/Qwen3.8-27B model from Alibaba, released on August 14 this year, is comparable to Anthropic Opus 4.6 released February 5 in multiple benchmarks, particularly coding and agentic tasks. People are reporting inference speeds of 200+ tokens per second which equates to 500 million tokens/month running on a PC equipped with a single RTX 5090 GPU https://www.nvidia.com/en-au/geforce/graphics-cards/50-series/rtx-5090/ . Tech Youtuber Alex Ziskind did the math https://x.com/digitalix/status/2091491916625875163 for a different model, Deepseek v4 Flash, a larger model running on a bigger machine: a DGX Station GB300: 17 billion tokens a month for $100,000USD is a reasonable business investment when you have a clear path to recouping your costs and the ability to serve an entire team for under $3kUSD/month while keeping complete control of your data and IP. But with the token price difference between closed models like Claude Opus and comparable open source models running locally, even a sub-$10k machine with an RTX 5090 can currently pay for itself if you can keep it busy. This may change in the future if neoclouds and the hyperscalars step up to price match or even beat prices for these models. Cerebras has already announced that Qwen3.8 27B is coming to its high speed inference platform https://x.com/cerebras/status/2088419572416147582 , but pricing hasn’t been announced. Will the future be lock-in or local? Driven by improvements across models, hardware, and inference software, “good enough” AI is dispersing out of the labs and into general availability. Businesses can tailor their model usage through routing services and handle as much of it on premises as they require, with AWS, Cloudflare and Google, as well as neoclouds, providing dedicated inference if they need it. This leaves management of the delicate dance around lock-in to the tools and processes that run on top of the inference layer – the skills, the workflows, the integration with your systems. We hope to look at this area in a future article. Until then, if want to chat about building software with experienced and now AI-assisted developers, drop us a line and let’s talk /contact-us/ .