Microsoft tests Kimi K3 for Copilot in bid to cut AI costs by $600 million Microsoft is testing Moonshot AI's Kimi K3 model for Copilot, aiming to reduce AI inference costs by up to $600 million by shifting workloads from OpenAI and Anthropic models. The Chinese startup's 2.8 trillion parameter model, released July 16, is being added to Azure for evaluation on response quality, reliability, safety, and latency. A successful deployment would deepen Microsoft's use of Chinese open-weight models and increase pricing pressure on US AI developers. Microsoft tests Kimi K3 for Copilot in bid to cut AI costs by $600 million Microsoft is adding Moonshot AI’s Kimi K3 to Azure and evaluating whether it can power Copilot features currently handled by OpenAI and Anthropic models. Microsoft is preparing to test Moonshot AI’s Kimi K3 model inside Copilot as the company looks to reduce its reliance on more expensive models from OpenAI and Anthropic. Microsoft is in the process of adding Kimi K3 to its Azure cloud service, while engineers working on Copilot plan to evaluate whether the model can power features that currently run on OpenAI and Anthropic systems, according https://www.theinformation.com/newsletters/ai-agenda/new-kimi-k3-model-means-u-s-china-ai-race to The Information. The potential shift could reduce Microsoft’s AI inference costs by as much as $600 million, according to the report. Microsoft has not publicly confirmed the estimate or detailed which Copilot features could move to Kimi K3. The company has previously tested DeepSeek models and earlier versions of Kimi for Copilot, according to a person familiar with its plans cited by The Information. Those evaluations indicate Microsoft is exploring a broader mix of models rather than relying on a single provider across every Copilot workload. The tests do not mean Kimi K3 has already been deployed in Copilot. Microsoft engineers must first determine whether the model meets the product’s requirements for response quality, reliability, safety and latency. Inference refers to the computing used each time a model receives a prompt and generates a response. The costs can become substantial when an AI service operates across millions of users and processes large volumes of tokens. Moonshot released Kimi K3 on July 16. The Chinese AI startup describes it as a 2.8 trillion parameter model with native multimodal capabilities and a one million token context window, designed for coding, knowledge work and complex reasoning. Kimi K3’s open weight structure could give Microsoft more control over how the model is hosted and optimized on Azure. It could also allow the company to assign less demanding Copilot tasks to cheaper models while reserving more expensive systems for workloads requiring stronger reasoning. Microsoft has increasingly promoted this model selection approach through Foundry. The company describes the platform as model agnostic and encourages developers to select models based on capability, safety, latency and cost rather than sending every request to the most powerful available system. Microsoft’s public Foundry pricing page currently lists earlier Moonshot models, including Kimi K2 Thinking, Kimi K2.5 Thinking and Kimi K2.6 Thinking. Kimi K3 has not yet appeared on the page, supporting the report that its Azure integration remains in progress. A successful Copilot test would deepen Microsoft’s use of Chinese open weight models beyond offering them to Azure customers. It would place Kimi K3 directly inside one of Microsoft’s flagship AI products while increasing pricing pressure on the US model developers currently powering its features. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .