{"slug": "at-t-cuts-ai-coding-costs-56-with-minimal-performance-decline", "title": "AT&T cuts AI coding costs 56% with minimal performance decline", "summary": "AT&T reduced its AI coding costs by 56% with only a 2% decline in performance by routing routine employee queries to cheaper open-source models via LiteLLM, processing 45 billion tokens daily on its internal 'Ask AT&T' platform. AT&T VP Mark Austin said open-source models are narrowing the capability gap with proprietary ones to 6-10 months, and the company aims to increase open-model usage from 40% to 60-70% of queries. Custom telecom-tuned models delivered up to 90% savings in inference costs at scale, and Goldman Sachs sees this as a template for enterprise AI cost management.", "body_md": "Via businesssearch.org\n\n# AT&T cuts AI coding costs 56% with minimal performance decline\n\nThe telecom giant is routing routine AI queries to cheaper open-source models, processing 45 billion tokens daily while barely denting output quality.\n\nAT&T found a way to slash its AI coding costs by more than half, and the trick is almost disappointingly simple: stop using the expensive model when a cheaper one works just as well.\n\nThe telecom giant implemented model routing technology through LiteLLM that redirects routine employee queries, particularly coding-related ones, toward lower-cost open-source models. The result was a 56% reduction in AI coding costs with only a 2% decline in performance quality. For a company processing roughly 45 billion tokens daily through its internal “Ask AT&T” platform, those savings add up fast.\n\n## The routing playbook\n\nThe concept behind AT&T’s approach is what the industry calls model routing, essentially a traffic cop for AI queries. Simple questions get sent to lightweight, inexpensive models. Complex tasks still go to premium options from OpenAI and Anthropic.\n\nAT&T VP Mark Austin noted that open-source models are narrowing the performance gap with their proprietary counterparts, with a difference of only 6-10 months in capabilities.\n\nCurrently, open models handle about 40% of employee AI queries at AT&T. The company is targeting 60-70% in the near term. The models doing the heavy lifting on the open-source side include Nvidia’s Nemotron, Meta’s Llama, and Google’s Gemma, with AT&T actively evaluating additional alternatives.\n\n## Telecom-tuned models push savings even further\n\nAT&T didn’t stop at generic model routing. Between February and July 2026, the company experimented with custom telecom-tuned models designed for industry-specific tasks. Those experiments delivered up to 90% savings in inference costs at scale.\n\nThe telecom-specific models are purpose-built for the kinds of queries AT&T employees actually make: network troubleshooting, customer service scripts, internal documentation lookups.\n\n## What this means for enterprise AI spending\n\nGoldman Sachs has flagged this trend as potentially advantageous for Big Tech firms, suggesting AT&T’s task-specific routing approach could serve as a template for cost management in enterprise AI deployment.\n\nFor other enterprises considering a similar move, the 45-billion-token-per-day figure is instructive. AT&T isn’t running a small pilot. This is production-scale deployment across a workforce of roughly 150,000 employees.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/at-t-cuts-ai-coding-costs-56-with-minimal-performance-decline", "canonical_source": "https://cryptobriefing.com/att-cuts-ai-coding-costs-model-routing/", "published_at": "2026-08-21 20:11:11+00:00", "updated_at": "2026-08-21 20:43:53.599312+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-products"], "entities": ["AT&T", "LiteLLM", "OpenAI", "Anthropic", "Nvidia", "Meta", "Google", "Goldman Sachs"], "alternates": {"html": "https://wpnews.pro/news/at-t-cuts-ai-coding-costs-56-with-minimal-performance-decline", "markdown": "https://wpnews.pro/news/at-t-cuts-ai-coding-costs-56-with-minimal-performance-decline.md", "text": "https://wpnews.pro/news/at-t-cuts-ai-coding-costs-56-with-minimal-performance-decline.txt", "jsonld": "https://wpnews.pro/news/at-t-cuts-ai-coding-costs-56-with-minimal-performance-decline.jsonld"}}