{"slug": "cloud-ops-is-different-in-a-neocloud", "title": "Cloud ops is different in a neocloud", "summary": "Enterprises are increasingly adopting neoclouds—specialized cloud providers such as CoreWeave, Lambda, and Crusoe Cloud built around AI infrastructure—for GPU-heavy workloads, drawn by better economics and faster access to capacity than traditional hyperscalers. However, maintaining these environments differs significantly from AWS, Azure, or Google Cloud, particularly in security, performance, and business continuity, requiring more direct enterprise ownership of security controls and processes.", "body_md": "Enterprises are taking a serious look at neoclouds, the specialized cloud providers built primarily around AI infrastructure, especially GPUs, high-speed networking, and large-scale compute clusters for model training and inference. Unlike traditional hyperscalers that provide broad platforms for almost every kind of enterprise workload, neoclouds tend to focus more narrowly on accelerated computing. CoreWeave, Lambda, Crusoe Cloud, and others are all commonly associated with this emerging AI infrastructure market.\n\nThe interest is not difficult to understand. Enterprises are under pressure to move [generative AI](https://www.infoworld.com/article/2338115/what-is-generative-ai-artificial-intelligence-that-creates.html), [machine learning](https://www.infoworld.com/article/2338768/the-engines-of-ai-machine-learning-algorithms-explained.html), and [advanced analytics](https://www.cio.com/article/228901/what-is-predictive-analytics-transforming-data-into-future-insights.html) projects out of the lab and into production. At the same time, access to large blocks of GPU capacity has become expensive, constrained, and in some cases difficult to obtain from the major hyperscalers. Many enterprises are finding that neoclouds can offer better economics, faster access to capacity, or configurations more closely aligned with AI workloads.\n\nThis does not mean AWS, Microsoft Azure, and Google Cloud are being displaced. They remain the default operating environment for most enterprise cloud deployments. They provide mature administrative planes, security tools, compliance frameworks, global footprints, managed services, and operational ecosystems that enterprises have spent years learning how to use.\n\nHowever, AI has changed the infrastructure conversation. Enterprises are worrying less about which cloud they are standardized on and focusing instead on getting the AI capacity they need, when they need it, at a price that does not destroy the business case.\n\nThat brings up one of the most common questions I get from clients: “How different is it to maintain these remote AI cloud systems compared with what we already do on AWS, Azure, or Google Cloud?” My answer is that the fundamentals of cloud operations still apply, but the administrative model does change in important ways. Neoclouds are not simply cheaper hyperscalers. They are specialized infrastructure environments, and specialization always creates trade-offs.\n\nThe biggest administrative differences show up in three areas: security, performance, and business continuity/disaster recovery.\n\nSecurity in the hyperscaler world is mature because the administrative ecosystem is mature. AWS, Microsoft, and Google have spent years building deeply integrated [identity systems](https://www.csoonline.com/article/518296/what-is-iam-identity-and-access-management-explained.html), key management services, logging tools, policy engines, compliance programs, network controls, vulnerability management capabilities, and security monitoring services. Enterprises still misconfigure these services all the time, but the building blocks are well known and widely understood.\n\nWith neoclouds, security administration may require more direct enterprise ownership. Some providers have strong security capabilities and mature operational practices. Others are still building out the kinds of enterprise-grade controls large organizations expect from the hyperscalers. That means administrators cannot assume that identity federation, privileged access controls, audit logging, encryption, network segmentation, and compliance reporting will behave in familiar ways.\n\nThis matters because AI workloads often involve some of the most valuable data an enterprise owns. Training sets, fine-tuning data, prompts, embeddings, model weights, vector databases, and inference outputs may contain intellectual property, customer data, regulated information, or confidential business logic. If an enterprise is using proprietary operational data to fine-tune a model, the administrative stakes are higher than simply spinning up remote compute.\n\nThe shared responsibility model still applies, but it must be examined provider by provider. Enterprises need to understand who controls encryption keys, how administrative access is granted and revoked, how logs are exported to the security operations center, how data is isolated between tenants, and how provider personnel access is governed. These are not paperwork questions. They are operating model questions.\n\nThe second difference is performance. Traditional cloud administration has trained enterprises to think in abstractions. Administrators select instance types, storage classes, managed databases, autoscaling policies, and observability dashboards. The underlying hardware matters, but it is usually hidden behind a service model.\n\nAI changes that. With neoclouds, performance administration often gets much closer to the physical infrastructure. GPU type, GPU memory, interconnect design, storage throughput, cluster topology, job scheduling, data locality, and network latency can all have a direct effect on whether an AI workload performs well or wastes money.\n\nGPU economics are unforgiving. An idle or underutilized GPU is a major financial problem. If data pipelines cannot feed accelerators fast enough, if distributed training is misconfigured, or if storage throughput becomes the bottleneck, the enterprise can quickly lose the cost advantage that made the neocloud attractive in the first place.\n\nAdministrators therefore need to understand more than basic cloud operations. They need to know how AI workloads behave at scale. They need to understand how training jobs consume storage and network resources, how inference demand fluctuates, how clusters are allocated, and how to measure actual accelerator utilization. This requires closer collaboration among cloud operations, AI engineering, data engineering, platform engineering, and finance.\n\nCapacity planning also changes. Hyperscalers created the expectation of near-infinite elasticity, even though that expectation has always been somewhat exaggerated. In the AI market, it is even less reliable. Neoclouds may provide better access to GPU capacity, but that capacity may come through reservations, fixed clusters, specific hardware commitments, or contractual windows. Administrators need to align training schedules, experimentation cycles, inference growth, and budget controls with the provider’s actual capacity model.\n\nPerformance administration in neoclouds is not just about watching dashboards. It is about managing workload economics at the infrastructure level.\n\nThe third difference is business continuity and disaster recovery. Too many enterprises still believe that if something runs in the cloud, resilience is included. That assumption is dangerous in any cloud environment, but even more so when dealing with specialized AI infrastructure.\n\nThe hyperscalers provide large global footprints, multiple regions, availability zones, replication services, backup tools, managed failover options, and well-documented resilience patterns. Neoclouds may not offer the same geographic depth or the same range of native continuity services. Administrators must be much more explicit about recovery objectives, failover design, replication, and restoration procedures.\n\nAI workloads complicate this further. Recovering an AI system is not the same as restoring a traditional application server. Enterprises need to protect data sets, training checkpoints, model artifacts, feature stores, [vector databases](https://www.infoworld.com/article/2335281/vector-databases-in-llms-and-search.html), orchestration pipelines, container images, configuration files, and inference endpoints. If a neocloud environment becomes unavailable, can the business restart training from a checkpoint? Can inference move to another environment? Can the same model run on different accelerators, drivers, frameworks, and networking assumptions?\n\nThose questions need answers before the outage, not during it. Some AI workloads can tolerate delay. A training job may be paused and restarted later without major business impact. Other workloads, especially production inference systems embedded in customer-facing processes, may require much more aggressive recovery targets.\n\nEnterprises should evaluate neoclouds with realistic expectations. The economics may open the door, and the capacity may make the decision urgent. The long-term success of neocloud adoption, however, will depend on how well enterprises administer the differences.", "url": "https://wpnews.pro/news/cloud-ops-is-different-in-a-neocloud", "canonical_source": "https://www.infoworld.com/article/4209350/cloud-ops-is-different-in-a-neocloud.html", "published_at": "2026-08-14 09:00:00+00:00", "updated_at": "2026-08-14 09:07:22.547585+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-policy"], "entities": ["CoreWeave", "Lambda", "Crusoe Cloud", "AWS", "Microsoft Azure", "Google Cloud"], "alternates": {"html": "https://wpnews.pro/news/cloud-ops-is-different-in-a-neocloud", "markdown": "https://wpnews.pro/news/cloud-ops-is-different-in-a-neocloud.md", "text": "https://wpnews.pro/news/cloud-ops-is-different-in-a-neocloud.txt", "jsonld": "https://wpnews.pro/news/cloud-ops-is-different-in-a-neocloud.jsonld"}}