Anyscale on Azure is Generally Available: Enabling Enterprises to Own the Full AI Loop, Not Just Inference Anyscale on Azure became generally available on October 7, 2026, letting enterprises run the full AI loop — data processing, training, post-training, evaluation, and inference — on Azure Kubernetes Service in their own subscription, with spend counting toward their existing Microsoft Azure Consumption Commitment. Anyscale said the platform runs on open source Ray, governed by the PyTorch Foundation, and applies customers' existing Azure RBAC and Azure Policy controls unchanged through their Microsoft Entra tenant. Microsoft Chairman and CEO Satya Nadella said it is "imperative that we distribute the learning infrastructure to every firm so that they can control their own learning loop. Anyscale on Azure is Generally Available: Enabling Enterprises to Own the Full AI Loop, Not Just Inference Katarina Stanley https://anyscale.com/blog?author=katarina-stanley and Adhip Gupta https://anyscale.com/blog?author=adhip-gupta | October 7, 2026 Today, Anyscale on Azure is generally available , so enterprises can run their own AI loops—spanning data processing, training, post-training, evaluation, and inference—with the ease of use they expect from hosted model inference. Renting intelligence by the token is the fastest way to prove an AI use case. AI-native companies built this way first, assembling products at speed on frontier model APIs, and the same APIs carried them to scale. But along the way, they gave up control over three things: the model, their data, and costs. The model and its roadmap belonged to someone else. The data that made a use case worth building had to leave the company’s environment to reach the model. And the bill rose in a straight line with usage, so scale never lowered with unit cost. Anyone who rents makes the same trade. AI-native companies ran into these limits first. Once a use case proved its value, cost and data exposure became the problem. To mitigate this, they moved to smaller, domain-specific models customized on their own data and run on their own GPUs, which gave them faster responses at lower cost. The bigger advantage came from improving those models continuously: owning the learning loop https://www.anyscale.com/blog/learning-loops that turns production data into training and post-training runs, and that turns the updated models into inference deployments that power their AI applications. That loop replaces an inference-only pipeline with a set of AI workloads – data processing, training and fine-tuning, post-training, and inference – that need to share GPU clusters. To run many workloads across many GPUs, Ray became the common runtime and orchestration layer for AI-native teams. Enterprises want the same path, extending the Azure governance they already trust to every step of the loop. Anyscale turns Ray from a distributed compute engine into an enterprise-ready platform: teams build and scale every step of the loop while Anyscale handles the cluster operations, all within the security and governance they've already built on Azure. Azure customers deploy it from the Azure portal, run production AI on Azure Kubernetes Service AKS in their own subscription, and pay for it like any other Azure service, with spend counting toward their existing Microsoft Azure Consumption Commitment MACC . With Anyscale on Azure, you can: - Own your AI. Anyscale runs on open source Ray, governed by the PyTorch Foundation. Your data pipelines, training code, and serving logic run on open source, across every layer of the stack https://www.anyscale.com/blog/ai-compute-open-source-stack-kubernetes-ray-pytorch-vllm . - Keep your data, evals, and models secure. Your workloads run on AKS in your own Azure subscription, associated with your organization's Microsoft Entra tenant. Your platform team's existing Azure RBAC and Azure Policy controls apply unchanged. - Increase output per GPU to reduce costs. Shared GPU capacity brings efficiencies of scale, and post-training on your own data lets teams move to smaller, domain-specific models that cost less to serve. LinkOwn your AI Satya Nadella made the case https://x.com/satyanadella/status/2076323181154230284 earlier this year: “It’s imperative that we distribute the learning infrastructure to every firm so that they can control their own learning loop… a company should be able to use a model without giving up the knowledge that makes it unique.” – Satya Nadella, Chairman and CEO, Microsoft Owning your AI means owning the loop that improves it. Production results feed data curation, curated data feeds training, and the next model version goes back into serving. Anyscale on Azure runs that whole loop on one platform. It enables you to curate your multimodal text, images, video, documents, etc data, scale model fine-tuning and other post-training techniques such as reinforcement learning RL , supports post-training and fine-tuning, and keeps models serving through updates and demand swings. For each of those, Anyscale orchestrates compute to match the demands of each workload. For example, to deploy LLMs and other models as endpoints, Anyscale Services https://docs.anyscale.com/services deploy with zero-downtime upgrades, scale to zero when idle, and compact replicas to keep GPU utilization high under real traffic. Every stage of that loop runs on Ray, which scales Python code. That means it works with the open-source AI frameworks your teams already use, like PyTorch and vLLM, and with new ones as they come out. Your pipelines and applications stay portable, because it's fully compatible with open-source APIs. Your teams stay flexible, because they can adopt new frameworks without changing how they scale. Because one engine runs every AI workload, teams can benefit in two ways. They can streamline their use of advanced techniques like reinforcement learning RL https://www.youtube.com/watch?v=QnxcDsMtaK8&pp=ygUYd3loeSBybCBpcyBoYXJkIGFueXNjYWxl which combine inference and training in one loop. At Ray Summit 2026, Microsoft AI shared that MAI-Thinking-1, its frontier reasoning model, was trained using Ray watch the session here https://www.youtube.com/watch?v=7fCwq7pIrkA . Or they can start with one workload and extend to the next without friction. BMW began using Ray for model training https://www.youtube.com/watch?v=s3x3STd5vuU , and when it was time to serve LLMs https://youtu.be/yKt-CpfrOIE?si=JC1nlZTBK0GtB6 a&t=35 , it integrated it to their AI Gateway to run inference for open-source models across the enterprise. LinkKeep your data, evals, and models secure Anyscale on Azure makes Ray production ready, scalable, and reliable in your own Azure tenant, with support from the engineers who build Ray. This means your data, evals, models, and container images stay in your subscription. As sovereignty and differentiation become central to AI strategy, this will be a requirement for every company, not just those with strict regulatory requirements. The network model is egress-only. Every connection originates from the Anyscale operator and Ray clusters in your AKS cluster, outbound to the Anyscale control plane, with no inbound firewall rules to open. The full deployment architecture, from the AKS data plane in your subscription to the Anyscale-hosted control plane, is diagrammed in the preview announcement. LinkReduce costs A per-token API charges the same rate whether you run one workload or fifty. On your own GPUs, every increase in utilization lowers the cost of each unit of work. Average GPU utilization on traditional ML infrastructure sits below 30% across the typical enterprise. That waste starts inside a single workload and compounds across teams, and Anyscale addresses it at every level. By optimizing scheduling and orchestration for individual workloads and across teams, Anyscale delivers 80%+ average GPU utilization. That means 2-3x more work from the same GPU fleet. At the workload level, GPUs often sit idle waiting for data: a handful of CPUs can't load and preprocess the next batch fast enough to keep them busy. Anyscale fixes this by spinning up as many CPU workers as needed to load data in parallel and stream it straight to the GPUs, so the next batch is ready the moment the GPU finishes the last one. Jobs also tend to reserve whole GPUs they only partly use, so fractional GPUs let you pack more workloads or models onto each device. Autoscaling adjusts resources before a job runs out of memory, and checkpointing restarts a failed job from its last completed step instead of from scratch, so a spot eviction or a straggler GPU no longer burns paid-for capacity on reruns. Across teams, the waste comes from siloed clusters. Each team reserves enough capacity for its peak, uses it only part of the time, and leaves it sitting idle while other teams wait for compute they can't access.. Anyscale pools those clusters into shared compute managed from a single pane, drawing on spot, reserved, and on-demand capacity as it becomes available. Scheduling is priority-aware, so queues follow business priorities, and when an urgent job lands, lower-priority work is paused rather than killed, then picks up where it left off. Across clusters and regions, that pool behaves as one. Demand for GPUs outpaces supply everywhere, so the teams running the largest AI workloads plan capacity as a portfolio. The Anyscale scheduler places workloads across clusters and Azure regions, with Azure as the control point for identity and policy. The same model extends to GPU capacity a customer attaches beyond Azure, with Azure remaining the home for identity, governance, and management. GPU utilization is one lever. The other is the model itself. Post-training and fine-tuning open models on your own data lets teams replace large general models with smaller, task-specific ones that cost less to serve. But the true cost includes building the model too, and running data processing, and post-training on separate infrastructure adds operational overhead that can wipe out the inference savings. Anyscale runs every step on one platform, so a smaller model doesn't mean a new stack to build or maintain for different steps of the AI lifecycle. LinkWhat customers are doing with Anyscale on Azure Xoople processes planetary-scale satellite imagery into decision-ready Earth intelligence on it. Hear about their deployment of Anyscale on Azure from Xoople’sVP of Engineering. LinkGetting started with Anyscale on Azure Provision Anyscale from the Azure portal and run your first Ray workload on AKS with your existing identity and networking. The deployment quickstarts on Microsoft Learn https://learn.microsoft.com/azure/anyscale-on-azure cover end-to-end setup. Anyscale Agent Skills https://www.anyscale.com/blog/announcing-anyscale-agent-skills-ray give your coding agent the knowledge to deploy, run, debug, and tune Ray workloads on Anyscale. LinkResources LinkGet connected - Learn more. http://anyscale.com/product/azure See the Anyscale on Azure product webpage. - Get started. https://docs.anyscale.com/clouds/azure Review the Anyscale on AKS documentation. - Talk to us . https://www.anyscale.com/contact-sales?utm-source=blog Request a demo or FDE support.