CoreWeave targets AI inference bottlenecks with full-stack optimization CoreWeave Inc. launched CoreWeave Forge, a platform connecting serving, observability, post-training and evaluation, and previewed CoreWeave RL Rollouts, which improved model reload latency by 15x versus a baseline configuration, according to Urvashi Chowdhary, CoreWeave's vice president of product and AI services. RL Rollouts is built on Nvidia Corp.'s Dynamo framework and loads new checkpoints into a live deployment to ease inference bottlenecks during reinforcement learning rollouts. A theCUBE Research survey of CoreWeave customers and prospects found one healthcare customer's inference workload share rose from about 10% in the first year to 40% in the second, with roughly 50% expected within 12 months. CoreWeave targets AI inference bottlenecks with full-stack optimization AI inference is fast becoming the workload that decides the economics of the AI boom. Training built the first wave of GPU clouds, but serving models faster and cheaper will define the next. That shift is pushing specialized cloud providers beyond raw GPU capacity https://siliconangle.com/2026/09/25/coreweave-full-stack-agentic-ai-fullyconnected/ into storage, networking and software. One provider is layering managed services for training, post-training and inference atop its infrastructure, according to Urvashi Chowdhary https://www.linkedin.com/in/urvashichowdhary/ pictured , vice president of product and AI services at CoreWeave Inc. “I think if you look at the AI developer’s journey, they’re looking to solve a problem and they want to do it as quickly as they can with the best performance and the cost to scale,” Chowdhary said. “So, we’ve been really focused on building up the layers of our stack, building on top of the reliable infrastructure to build more managed services, whether it’s for training, post-training, inference.” Chowdhary spoke with theCUBE Research’s Dave Vellante https://www.linkedin.com/in/dvellante/ and John Furrier https://www.linkedin.com/in/furrier/ at the Fully Connected event https://www.thecube.net/events/coreweave/coreweave-fully-connected-2026 , during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed AI inference, managed services across the full stack and CoreWeave RL Rollouts, a new capability designed to accelerate agentic model iteration. Disclosure below. Optimizing the AI inference stack layer by layer Inference demand is climbing fast. A survey of CoreWeave customers https://thecuberesearch.com/327-breaking-analysis-coreweaves-next-test-from-gpu-scarcity-to-a-durable-ai-cloud/ and prospects by theCUBE Research found one healthcare customer’s inference workload share rose from about 10% in the first year to 40% in the second, with a roughly 50% share expected within 12 months. CoreWeave’s answer is tuning every layer above the hardware, from the vLLM engine to quantized models and custom speculative decoders, Chowdhary explained. “One thing that we’ve been very intentional about is leveraging open source tools and technologies, contributing back to open systems so customers have flexibility and then also building our services on top of each other,” she said. Reinforcement learning adds new pressure. When customers train agentic models with rewards and verifiers, inference becomes the bottleneck during rollouts, according to Chowdhary. CoreWeave RL Rollouts https://www.coreweave.com/news/coreweave-forge-launches-turning-the-ai-loop-production-run-into-a-better-model-and-agent , a preview capability built on Nvidia Corp.’s Dynamo framework https://blogs.nvidia.com/blog/coreweave-agentic-ai-vera-rubin/ , loads new checkpoints into a live deployment. In testing, the capability improved model reload latency by 15x compared with a baseline configuration. “When you’re doing RL rollouts, you’re continually creating new model checkpoints and versions and you want those to roll out into your inference setup so you can scale it independently,” she said. “And we were able to speed that up by 15x, which means your training runs fast and your inference is scaling while you’re continuing to train your model quickly.” Those capabilities now sit inside CoreWeave Forge https://www.coreweave.com/blog/coreweave-forge-turn-ai-iteration-into-compounding-improvement , a platform launched at the event that connects serving, observability, post-training and evaluation. Forge is free to start, with paid tiers offering additional capabilities, extending access to AI development tools and services to individual developers, Chowdhary explained. “We want to create accessibility for these leading technologies,” she said. “Even if you’re an individual developer signing up today on your own, you still get the best performance, you still get the best reliability and you don’t have to compromise.” Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the Fully Connected event https://www.thecube.net/events/coreweave/coreweave-fully-connected-2026 : Disclosure: TheCUBE is a paid media partner for the the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE. Photo: SiliconANGLE A message from John Furrier, co-founder of SiliconANGLE: Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network , where technology leaders connect, share intelligence and create opportunities. - 15M+ viewers of theCUBE videos , powering conversations across AI, cloud, cybersecurity and more - 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/ https://siliconangle.com/aws-marketplace/ About SiliconANGLE Media SiliconANGLE https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552 , theCUBE Network https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da , theCUBE Research https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f , CUBE365 https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6 , theCUBE AI https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683 and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI. Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.