Cisco and Nvidia take AI factories from rack to runtime
AI factories are moving from ambitious plans toward production, but the path from graphics processing unit acquisition to usable systems remains a race against time.
Neoclouds already have customers waiting for capacity, enterprises are looking to bring inference workloads closer to home and sovereign AI programs are being built now. Those distinct buyer motions are converging around a common need: getting AI systems into production quickly enough to support the business, according to Will Eatherton (pictured, left), senior vice president of Cisco Systems Inc.
“Enterprises, many of them … are spending a large amount on tokens right now,” he said. “The rush and the pressure is getting these systems up so they can start off what has been an [application programming interface] into using local inference. I think it’s all converging on common architectures [and] common systems, but I think speed is either what’s broken or the challenge.”
Eatherton, along with Gilad Shainer (center), senior vice president of networking at Nvidia Corp., and Marc Hamilton (right), vice president of solutions architecture and engineering at Nvidia, spoke with theCUBE Research’s John Furrier at the “Cisco Secure AI Factory With Nvidia Expands to Rack Scale” event during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the rise of AI factories, rack-scale compute, Nvidia’s reference architecture, networking and the challenge of moving from the first token to sustained operations. ( Disclosure below.)*
AI factories move toward a unified deployment model
Cisco is expanding its Secure AI Factory beyond its networking foundation to include full rack-scale compute. The offering brings liquid-cooled systems into a broader solution, according to Eatherton.
“What we’re announcing today is that across the sovereign, enterprise and neocloud markets, we need to go big,” he said. “We are going broader with compute: We have partnered with Supermicro, bringing in the full rack scale. That is liquid-cooled, starting with Blackwell and moving to Vera Rubin. That gives us a breadth of compute systems that we can then wrap around from a Cisco standpoint, from a sales support and a software standpoint.”
To make the rack-scale expansion deployable, Cisco and Nvidia are building around Cisco Validated Designs that comply with the Nvidia Cloud Partner reference architecture. The framework is intended to reduce late-stage integration problems as customers bring complex AI systems into production, according to Hamilton.
“Traditional enterprises had server teams and networking teams … and the two didn’t come together until very late,” he said. “In an AI factory, because it’s a five-layer cake, everything has to work together. There are so many mistakes when customers try to go to one vendor and order networking [and] another vendor to order servers.”
Getting an AI factory to its first token is only the beginning. Cisco is also focused on monitoring, availability, software upgrades and lifecycle management as customers operate these systems over time, according to Eatherton.
“A lot of the industry focus is up to the point that you light up the cluster and you get your first token out,” he said. “That’s been a big focus. But the day-two aspects around monitoring, health, availability and software upgrades … are things that, from a Cisco standpoint, we’ve put a lot of focus on here over the years.”
The partnership also leaves room for flexibility within the networking stack. For AI factories, that flexibility gives customers room to customize the technology without abandoning familiar operating tools. The open interfaces in Nvidia Spectrum-X allow customers to run their own technologies on top of the platform, while Spectrum-X licensing allows Cisco Silicon One switches to connect to the access network, according to Shainer.
“That’s the reason that we build the reference architecture: to make sure that [customers] can get the full performance of Nvidia components,” he said. “That guarantee is now coming with the Cisco AI Factory. That’s how both of us guarantee that they’re getting the best performance out of their investment.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the “Cisco Secure AI Factory With Nvidia Expands to Rack Scale” event:
( Disclosure: TheCUBE is a paid media partner for the “Cisco Secure AI Factory With Nvidia Expands to Rack Scale” event. Neither Cisco, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)*
Photo: SiliconANGLE
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.