AT&T slashes Anthropic costs by up to 90% with aggressive open source AI pivot AT&T has shifted up to 25% of its 45 billion daily AI tokens to open-source models, cutting costs by 80% to 90% on targeted applications and saving millions, with plans to route 70-80% of workloads through open models. Andy Markus, AT&T's Chief Data and AI Officer, said the move is a strategic choice, supported by a proprietary cache-aware AI Gateway and the launch of OTel 2.0, trained on over 400 billion tokens using AMD GPUs and Microsoft Foundry. Via businesssearch.org AT&T slashes Anthropic costs by up to 90% with aggressive open source AI pivot The telecom giant processes 45 billion AI tokens daily and plans to route up to 80% of workloads through open models, saving millions in the process. AT&T has begun shifting significant portions of its artificial intelligence workloads away from Anthropic’s proprietary models toward open-source and open-weight alternatives, achieving cost reductions of 80% to 90% on targeted applications. The company currently processes roughly 45 billion AI tokens per day. About 25% of that volume already runs on open models, primarily for functions like network management. AT&T’s goal is to push that number to somewhere between 70% and 80% over time. The math behind the migration AT&T says the shift has already saved millions of dollars, though the company hasn’t disclosed a precise total figure. The savings aren’t just coming from swapping one model for a cheaper one. AT&T has built a proprietary “cache-aware AI Gateway” that handles intelligent prompt routing, essentially directing each query to the most cost-effective model capable of handling it well. Andy Markus, AT&T’s Chief Data and AI Officer, has positioned this approach as a deliberate strategic choice rather than a compromise. AT&T recently launched OTel 2.0, a specialized model trained on more than 400 billion tokens. The training infrastructure relied on advanced AMD GPUs paired with Microsoft Foundry. The open model moment For AT&T specifically, the appeal goes beyond pure cost savings. Running open models on your own infrastructure means sensitive telecom data, including network configurations, customer usage patterns, and infrastructure details, never leaves your environment. AMD, which provided the GPUs for OTel 2.0’s training, and Microsoft, which contributed Foundry infrastructure, are positioned on the enabling side of this shift. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .