CoreWeave brings Nvidia Vera Rubin, AI tools to its cloud services CoreWeave announced at its Fully Connected conference in San Francisco that Nvidia's Vera Rubin NVL72 rack-scale AI platform is available in its cloud, with Cognition, developer of the Devin software-engineering agent, as the first customer after bringing up a Vera Rubin cluster in early September. Cognition reported up to a 4.8X increase in total token throughput for SWE-2 inference workloads on Vera Rubin NVL72 versus an Nvidia GB200 NVL72 baseline, plus a 3.8X gain in output-token throughput for reinforcement-learning workloads. CoreWeave also added support for the Nvidia Vera CPU in rack-scale configurations of 128 Vera CPUs (11,264 cores) supporting more than 11,000 concurrent agent environments, and launched Forge, an AI development platform unifying training, inference, evaluation and agent development. CoreWeave https://www.networkworld.com/article/4018621/coreweave-acquires-core-scientific-for-9b-to-power-ai-infrastructure-push.html has made a series of announcements starting with Nvidia’s new Vera Rubin https://www.networkworld.com/article/4188058/nvidia-unveils-vera-rubin-platform-targeting-ai-hpc-infrastructure-customers.html NVL72 rack-scale AI platform available in its cloud. At its CoreWeave’s Fully Connected conference in San Francisco, the company announced Cognition, the developer of the Devin AI software-engineering agent, is the first customer for a Vera Rubin cluster. Cognition brought up the Vera Rubin cluster https://www.engineering.com/coreweave-adds-nvidia-vera-cpu-to-its-compute-portfolio/ in early September and ran what CoreWeave described as the first customer-executed inference benchmark on the new platform. In announcing the deal, Cognition said it measured up to a 4.8X increase in total token throughput for SWE-2 inference workloads on Vera Rubin NVL72 compared with an Nvidia GB200 NVL72 baseline. The company also reported a 3.8X gain in output-token throughput for reinforcement-learning workloads. Chen Goldberg https://www.linkedin.com/in/goldbergchen/ , executive vice president of product and engineering at CoreWeave, said the company’s engineering investment was intended to allow customers to bring new rack-scale systems into production within days rather than rebuilding infrastructure processes for each GPU generation. “When it comes to agentic tasks, long contexts, repeated model calls and thousands of concurrent tasks put pressure on the entire platform,” Goldberg said in a statement. “Our job is to make compute, networking and software work as a single system.” He said the throughput gains could translate into more concurrent Devin sessions per GPU, faster research cycles and a lower per-session cost without reducing generation speed. CoreWeave has already deployed Nvidia GB200 and GB300 NVL72 systems https://www.networkworld.com/article/4125806/eying-ai-factories-nvidia-buys-bigger-stake-in-coreweave.html , including Dell PowerRack configurations developed with Dell Technologies. The cloud provider said its existing customers will be able to operate Vera Rubin capacity using the same tooling and operating approach used for their GB200 and GB300 fleets. In addition to the Vera Rubin support, CoreWeave is adding support for the Vera CPU, positioning the processor as a way to address the growing amount of CPU-intensive work surrounding agentic AI workloads. While GPUs do the heavy lifting for model training and inference, CPUs handle important tasks such as isolated execution environments, reinforcement-learning workloads, tool calls, code execution and data pipelines, the company says. Nvidia describes Vera as the first CPU designed specifically for AI agents and the workloads surrounding modern AI systems. CoreWeave’s initial Vera deployment is built around rack-scale configurations containing 128 Vera CPUs, or 11,264 CPU cores in a single rack. The rack also incorporates Nvidia BlueField-4 DPUs and Spectrum-X Ethernet switching to provide connectivity between Vera nodes. CoreWeave said the configuration provides enough CPU capacity to support more than 11,000 concurrent agent environments. On the software side, CoreWeave launched Forge https://www.coreweave.com/blog/coreweave-forge-turn-ai-iteration-into-compounding-improvement , a new AI development platform designed to connect model training, inference, evaluation and agent development in a single environment. The company said Forge is intended to address a growing challenge of development workflows are often spread across multiple tools and vendors, making it difficult to carry information from production systems back into training and subsequent model iterations. Forge brings together experiment tracking, evaluation, agent observability, post-training, inference and model management. The platform is designed to let teams use their preferred models, frameworks and cloud environments rather than locking them into a single AI stack, CoreWeave stated. The company described Forge as a continuous AI development loop in which teams can run models and agents, observe their behavior, curate data, make improvements, evaluate results and repeat the process. CoreWeave said it is positioning its software stack as a way for customers to move between Nvidia GPU generations without changing their operational model. The company’s platform includes its Kubernetes service, SUNK, Mission Control, Sandboxes and serverless-inference offerings. CoreWeave said Forge is available now, with a 30-day free trial offered for the Pro edition.