AT&T cuts AI coding costs 56% with minimal performance decline AT&T reduced its AI coding costs by 56% with only a 2% decline in performance by routing routine employee queries to cheaper open-source models via LiteLLM, processing 45 billion tokens daily on its internal 'Ask AT&T' platform. AT&T VP Mark Austin said open-source models are narrowing the capability gap with proprietary ones to 6-10 months, and the company aims to increase open-model usage from 40% to 60-70% of queries. Custom telecom-tuned models delivered up to 90% savings in inference costs at scale, and Goldman Sachs sees this as a template for enterprise AI cost management. Via businesssearch.org AT&T cuts AI coding costs 56% with minimal performance decline The telecom giant is routing routine AI queries to cheaper open-source models, processing 45 billion tokens daily while barely denting output quality. AT&T found a way to slash its AI coding costs by more than half, and the trick is almost disappointingly simple: stop using the expensive model when a cheaper one works just as well. The telecom giant implemented model routing technology through LiteLLM that redirects routine employee queries, particularly coding-related ones, toward lower-cost open-source models. The result was a 56% reduction in AI coding costs with only a 2% decline in performance quality. For a company processing roughly 45 billion tokens daily through its internal “Ask AT&T” platform, those savings add up fast. The routing playbook The concept behind AT&T’s approach is what the industry calls model routing, essentially a traffic cop for AI queries. Simple questions get sent to lightweight, inexpensive models. Complex tasks still go to premium options from OpenAI and Anthropic. AT&T VP Mark Austin noted that open-source models are narrowing the performance gap with their proprietary counterparts, with a difference of only 6-10 months in capabilities. Currently, open models handle about 40% of employee AI queries at AT&T. The company is targeting 60-70% in the near term. The models doing the heavy lifting on the open-source side include Nvidia’s Nemotron, Meta’s Llama, and Google’s Gemma, with AT&T actively evaluating additional alternatives. Telecom-tuned models push savings even further AT&T didn’t stop at generic model routing. Between February and July 2026, the company experimented with custom telecom-tuned models designed for industry-specific tasks. Those experiments delivered up to 90% savings in inference costs at scale. The telecom-specific models are purpose-built for the kinds of queries AT&T employees actually make: network troubleshooting, customer service scripts, internal documentation lookups. What this means for enterprise AI spending Goldman Sachs has flagged this trend as potentially advantageous for Big Tech firms, suggesting AT&T’s task-specific routing approach could serve as a template for cost management in enterprise AI deployment. For other enterprises considering a similar move, the 45-billion-token-per-day figure is instructive. AT&T isn’t running a small pilot. This is production-scale deployment across a workforce of roughly 150,000 employees. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .