cd /news/artificial-intelligence/at-t-routed-the-easy-tokens-off-the-… · home topics artificial-intelligence article
[ARTICLE · art-126081] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AT&T Routed the Easy Tokens Off the Frontier Bill

AT&T has shifted roughly 40 percent of its AI usage to open-weight models, up from 20 percent in May, according to chief data and AI officer Andy Markus, cutting costs by as much as 80 percent for high-volume tasks like transcription and customer service. The company customizes Google's Gemma and Meta's Llama rather than Chinese open models, while reserving frontier systems from Anthropic and OpenAI for harder coding and media generation work. The shift, echoed by Airbnb and Deloitte, suggests large buyers are moving easy tokens off expensive closed APIs.

by read5 min views3 publishedSep 10, 2026

AT&T used to buy the closed package for the obvious jobs. Customer service, call transcription, coding help. Anthropic and OpenAI. Fees, no weights. Eli Tan's New York Times story on September 4 is the update. Andy Markus, the company's chief data and AI officer, said open models were 20 percent of AT&T's AI use by May, 40 percent now, and maybe 60 percent in the coming months. Cost versus earlier this year is down as much as 80 percent. Airbnb and Deloitte show up in the same piece doing a version of the same move.

I take the 80 percent cut as a real number, not a slogan. A telco's transcription and ticket macros are high-volume work with a grade you can actually compute. You do not need the model that wins a reasoning contest to turn a call into text. If Markus can park that layer on a downloadable model and keep Claude or GPT for the leftover, the invoice moves even if the leaderboard does not.

Jerry Tang, who runs Atlas Cloud, gave Tan the split that matches that invoice. Closed models still look better on heavy coding and on image and video generation. Open models often win on simple, specialized jobs. AT&T is running that split in production. Markus said the company takes open-weight models and customizes them for transcription and customer service. Llama did not replace the frontier stack. The frontier stack was sitting on a lot of tokens it did not need to touch.

OpenRouter's figure will get quoted as if it were a market-share print. Tan wrote that last month open models accounted for 58 percent of AI use, up from 10 percent a year ago, according to U.S. user data from the routing platform. OpenRouter is a mall for people who already want to pick among models. It is a decent gauge of substitution among shoppers. It is a poor census of what a bank, a hospital, or a regulated telco spends on contracted APIs, on-prem GPUs, and Microsoft or Google bundles that never hit OpenRouter. Read 58 percent as evidence that the outside option is being used. Do not read it as 58 percent of corporate AI budgets walking out of the closed labs.

The geography of that outside option is the awkward part. Tan notes that many of the most popular open models come from Moonshot, DeepSeek, and Alibaba. Tang put some Chinese open models at 80 to 90 percent of closed-lab capability at as little as 20 percent of the price. AT&T is not taking that trade. Markus said the company researches Chinese models and does not use them. It is working with Google's Gemma and Meta's Llama instead. Last month Meta made its leading system an open-weights model and said it would price below other top U.S. models. Mark Zuckerberg's accompanying line was that blocking foreign open-source models is not an effective solution, and that American open-source models should be the best globally.

So the Fortune 500 version of getting hooked is narrower than the headline. The buyer wants cheaper tokens and the right to fine-tune. The buyer still wants a U.S. name on the weights. Meta and Google pick up distribution they could not get as closed APIs. Nvidia's $12.9 billion Hugging Face purchase, which Tan flags in the same story, sits on that stack. Someone has to host, rank, and operationalize the downloads. The inference bill does not vanish. It moves from a lab subscription onto servers, chips, and a catalog.

That movement is the listing problem. Tan's piece says OpenAI and Anthropic are heading toward large offerings and have to charge more because they are spending billions on research and computing power. Neither company commented. A mix shift of the AT&T shape does not require the frontier product to fail. It requires the easy tokens to stop subsidizing it. If 40 percent of a large buyer's volume leaves, and the rest is the hard work, the lab can still have a business. The business has to look more like scarce escalation capacity and less like a default meter on every employee prompt.

David Stout of WebAI told Tan that many open models now run on a phone or a laptop, while leading closed models are getting more complicated and more compute-hungry. That is a cost-curve split. It also explains why open does not mean free. Markus landed on models that can be downloaded and modified without payment or approval. AT&T still pays for the machines that run them, the people who customize them, and the evals that decide when a job is too hard for the cheap path. Dario Amodei's case for tighter control of frontier weights is a different argument. It does not stop a telco from routing call summaries off the expensive API.

Three ways this can go. For the closed labs, the kind case is that AT&T-style routing caps the cheap work and leaves them a high-willingness-to-pay remainder that can fund the next training run. The middle case is a fork. More large U.S. shops copy the 20-40-60 staircase on Gemma and Llama, OpenRouter shoppers keep using Chinese models, and the two markets stop being the same market. The ugly case for the listings is that the remainder shrinks too. Once transcription quality is close enough, coding assistants get the same treatment AT&T already started, and the premium slice is a thin set of multimodal and high-stakes jobs.

Watch Markus's next mix print, and watch whether the 80 percent cost cut shows up as a smaller Anthropic and OpenAI line or as a larger GPU line. Those two ledgers will not move together.

Originally published at https://deanlee.info/essays/att-open-model-mix/.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @at&t 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/at-t-routed-the-easy…] indexed:0 read:5min 2026-09-10 ·