AWS and NVIDIA to Deploy 2 Million More GPUs in 2027-2028, Bringing Vera CPUs and Custom NVHBM to Trainium Amazon Web Services and NVIDIA announced plans to deploy 2 million additional NVIDIA GPUs across AWS data centers between 2027 and 2028, expanding their 16-year collaboration to include NVIDIA Vera CPUs and custom NVHBM memory for Trainium processors. The multi-year rollout will integrate Blackwell Ultra, Rubin, and Rubin Ultra GPUs, and the companies will co-develop dedicated AI factories with 100,000 GPUs for federal and national security workloads. AWS CEO Matt Garman said customers want the freedom to choose the best tools for AI workloads, while NVIDIA CEO Jensen Huang said demand is running ahead of every forecast. Amazon Web Services and NVIDIA have announced a major expansion of their joint infrastructure initiatives, outlining plans to deploy an additional 2 million NVIDIA GPUs across AWS data centers between 2027 and 2028. Building on a 16-year engineering relationship, the expanded collaboration targets end-to-end scaling of AI infrastructure, encompassing accelerator capacity, host CPU architectures, high-bandwidth interconnects, sovereign government deployments, and accelerated data-processing pipelines. “Customers want the freedom to choose the best tools for their AI workloads, and they want confidence that everything works seamlessly together,” said Matt Garman, CEO of AWS. Jensen Huang, founder and CEO of NVIDIA, said the two companies “have built one of the great growth engines of the AI era, and demand is running ahead of every forecast,” pointing to an expansion that spans GPUs, CPUs, networking, open models, and software to make agentic and physical AI work at scale. Scaled GPU Deployments and EC2 G7 Instances The multi-year rollout will integrate 2 million NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs into the AWS Global Infrastructure, specifically designed to power large AI factories handling distributed training, scientific simulation, enterprise automation, and agentic workflows. AWS is also expanding its current-generation Blackwell fleet by introducing NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs in Amazon EC2 G7 instances. AWS is the first major cloud provider to offer instances accelerated by the RTX PRO 4500, which delivers up to 4.6 times higher AI inference performance and 2.1 times higher graphics throughput compared to previous-generation EC2 G6 instances. To maintain cluster-wide throughput and low latency across these high-density nodes, the companies are collaborating on NVIDIA Spectrum networking https://www.storagereview.com/news/nvidia-groq-3-lpx-enters-full-production-3400-tokens-per-second-at-100k-context-256-lp30s-per-rack optimizations tailored for large GPU training clusters. NVIDIA Vera CPUs and NVLink Fusion with Custom HBM To address the compute requirements of multi-step agentic AI workloads, AWS and NVIDIA are collaborating to bring NVIDIA Vera CPU https://www.storagereview.com/news/spacexai-adopts-nvidia-vera-cpus-for-grok-with-a-vera-rubin-nvl72-bound-for-orbit-in-starmind -based infrastructure to AWS. NVIDIA describes Vera as purpose-built for the next generation of AI, giving AWS customers a high-performance CPU option alongside accelerated infrastructure and complementing AWS’s strategy of offering the broadest possible choice of compute, from its own custom silicon to partner CPUs and accelerators. In the interconnect and memory domain, NVIDIA and Amazon’s Annapurna Labs are extending their partnership around NVIDIA NVLink Fusion. Originally introduced for next-generation AWS Trainium processors, NVLink Fusion is now being paired with NVIDIA’s custom high-bandwidth memory NVHBM technology in partnership with memory suppliers. NVIDIA says the combination would give Trainium access to faster, more power-efficient memory, allowing Annapurna Labs to combine Trainium silicon and NVIDIA GPUs within a shared, scale-up rack architecture. Dedicated AI Factories for Federal and National Security Workloads For government-sector requirements, AWS and NVIDIA are co-developing dedicated AI factories equipped with 100,000 GPUs, deployed across isolated, secure AWS infrastructure. The environment is engineered to support sensitive federal agency and defense workloads, providing compliance and operational security for classifications at Impact Level 6 IL6 and above. Core Platform Integrations: Nitro, Data Processing, and Robotics The expanded alliance continues to leverage existing core technical integrations across the AWS stack: AWS Nitro System and EFA: All NVIDIA GPU and Trainium-based EC2 instances rely on the AWS Nitro System for offloaded virtualization and hardware security, paired with Elastic Fabric Adapter EFA networking to deliver line-rate, low-latency scale-out interconnectivity. Managed Open Models: The NVIDIA Nemotron open model family remains natively integrated into AWS, offered as fully managed, serverless endpoints on Amazon Bedrock and as customizable foundation models on Amazon SageMaker. Hardware-Accelerated Analytics and Search: Amazon EMR utilizes Amazon EC2 G7 instances and the NVIDIA cuDF library to deliver up to 3.7 times faster data processing throughput and a 30 percent improvement in price-performance over standard CPU configurations. Additionally, Amazon OpenSearch Service leverages dedicated GPUs via NVIDIA CUDA-X cuVS libraries for vector indexing, achieving up to 9x faster index build times at approximately 1/4 the infrastructure cost across managed and serverless deployments. Physical AI and Robotics Automation: Amazon Robotics is incorporating NVIDIA’s physical AI portfolio, spanning NVIDIA Jetson edge compute modules, Omniverse simulation libraries, and the Isaac robotics platform. Running on GPU-accelerated EC2 instances, this pipeline supports large-scale synthetic data generation, simulation, robot training, route optimization, functional safety, and real-to-sim validation for automated warehouse systems.