cd /news/ai-infrastructure/wiring-and-powering-gpus-differently… · home › topics › ai-infrastructure › article
[ARTICLE · art-143840] src=siliconangle.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Wiring and powering GPUs differently can swing AI latency by orders of magnitude, says CoreWeave

CoreWeave Inc. senior vice president of AI initiatives Lukas Biewald said at the company's Fully Connected event that GPU networking and power distribution choices can swing AI latency by "orders of magnitude," as CoreWeave launched CoreWeave Forge, a development layer running training, inference, evaluation and agent development in one connected environment. LlamaIndex Inc. co-founder and CEO Jerry Liu said his company's workload is now about 75% inference and 25% training, processing millions of document pages per day for finance, legal and insurance customers while owning no GPU cluster. CoreWeave is expanding beyond GPU compute into networking, storage and software, and Biewald said it follows Nvidia Corp.'s standard networking protocols rather than the proprietary APIs used by hyperscalers such as Amazon Web Services Inc.

by read4 min views5 publishedOct 2, 2026
Wiring and powering GPUs differently can swing AI latency by orders of magnitude, says CoreWeave
Image: Siliconangle (auto-discovered)

Wiring and powering GPUs differently can swing AI latency by orders of magnitude, says CoreWeave

The neocloud market is moving past its origins as a stopgap for scarce graphics processing units. AI-native startups now choose their infrastructure on latency, burst capacity and openness, not just chip availability.

That shift is playing out at CoreWeave Inc., which is expanding beyond GPU compute into networking, storage and software as inference demand grows. At the same time, LlamaIndex Inc. has evolved from an open-source framework for retrieval-augmented generation into a model builder that rents its compute rather than owning it, according to Jerry Liu (pictured, right), co-founder and chief executive officer of LlamaIndex.

“We’re effectively a specialized AI lab right now that’s purely focused on building models for document parsing and extraction. We post-train open-weight models, we gather our own datasets and we make it really, really good at analyzing and reading documents to basically extract that data,” Liu said. “We care a lot about making sure that we can actually tailor everything we’re doing at the Pareto frontier of performance, cost, and latency for our customers.”

Liu and Lukas Biewald (left), senior vice president of AI initiatives at CoreWeave, spoke with theCUBE’s John Furrier and Dave Vellante at Fully Connected, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed long-running agents, governance and why AI-native startups are turning to specialized clouds for inference-heavy workloads. ( Disclosure below.)*

Why bursty AI workloads favor the neocloud model

LlamaIndex’s compute footprint barely existed a year ago. Today its workload runs about 75% inference and 25% training, and it processes millions of document pages per day for finance, legal and insurance customers whose paperwork arrives in bursts, according to Liu. The company owns no GPU cluster, so guaranteed capacity matters more than hardware ownership.

“We serve a lot of different customers at extremely persistent and also spiky workloads,” ,” Liu said. “We really, really need to make sure that we have the right capacity to serve our customers without getting throttled.”

CoreWeave is betting that capacity alone is not the differentiator. Biewald joined the company through its acquisition of Weights & Biases, the AI observability startup he co-founded, and CoreWeave used the event to launch CoreWeave Forge, a development layer that runs training, inference, evaluation and agent development in one connected environment. Coming from software, he initially questioned how much chip configuration could really matter, Biewald noted.

“I’ll tell you, the answer is ‘massive difference,'” Biewald said. “I’m talking orders of magnitude difference depending on how you do the networking for the chips [and] how you do the power distribution.”

Openness also separates CoreWeave from the hyperscalers, according to Biewald. Where providers such as Amazon Web Services Inc. lean on proprietary application programming interfaces that make workloads hard to move, CoreWeave follows the standard networking protocols recommended by Nvidia Corp., which brings broader open-source support. Analysts have observed that CoreWeave is broadening its portfolio much as AWS did in its early days, even as the company bristles at the neocloud label.

“CoreWeave knows that everyone is coming from a different cloud,” Biewald said. “Everyone’s going to host their web service on AWS or GCP, not on CoreWeave. CoreWeave is okay with that, so CoreWeave plays much more nicely with the other clouds.”

That ecosystem points to a larger change in who gets to build intelligence, Liu noted. Post-training a small open-weight model remains a skill limited to a narrow group of specialists today. Abundant neocloud capacity, combined with fast-improving coding agents, could open that work to far more people.

“Everyone is starting to get really good at defining observability and evals and the right metrics to focus on,” Liu said. “I think there’s going to be a world where we’re basically just going to automate this entire loop and make it accessible to everybody.”

Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of Fully Connected: ( Disclosure: TheCUBE is a paid media partner for the Fully Connected 2026 event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*

Photo: SiliconANGLE

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos , powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

About SiliconANGLE Media

SiliconANGLE,

theCUBE Network,

theCUBE Research,

CUBE365,

theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @coreweave inc. 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/wiring-and-powering-…] indexed:0 read:4min 2026-10-02 · —