Google reportedly developing ‘Frozen v2’ AI chip optimized for Gemini models
Google LLC is reportedly developing a new chip optimized to run its Gemini series of artificial intelligence models.
Sources told The Information today that the processor is codenamed “Frozen v2.” According to the publication, it’s expected to provide between six and 10 times better performance per watt than the search giant’s current silicon. Shares of Google parent Alphabet Inc. rose 1.5% on the report.
Off-the-shelf chips often contain components that customers don’t need. If an AI startup buys a graphics card that includes both inference and rendering cores, it may end up leaving the latter modules idle. That can lead to inefficiencies in AI projects.
Designing a custom chip addresses the challenge. A company can leave out circuits that aren’t needed for its workloads and thereby lower manufacturing costs. Alternatively, it can replace the unnecessary circuits with cores optimized for its use case.
Google has long offered custom AI chips called TPUs through its public cloud. The two newest entries in the lineup, the TPU 8t and TPUi, are optimized for training and inference, respectively. According to today’s report, Google plans to take that customization a step further by optimizing its upcoming Frozen v2 chip for its Gemini models’ architecture.
The significant efficiency gains that the company reportedly expects to unlock will be achieved through multiple routes. In particular, the company hopes that Frozen v2 will reduce the number of calculations needed to run Gemini. It also will reportedly reduce data movement.
Data movement is an issue when a neural network can’t fit into a graphics card’s onboard memory. When that happens, neural network elements are stored in off-chip storage. The graphics card must regularly move those elements from the off-chip storage to its logic circuits and vice versa, which slows down processing.
Google could address that bottleneck by equipping Frozen v2 with enough memory to run Gemini fully on-chip. Such a design would remove the need to move data to and from off-chip RAM.
The fact that Frozen v2 is expected to reduce the number of calculations needed to run Gemini suggests that it will also support a form of operator fusion. That’s a widely used technique for speeding up AI models. It works by combining several calculations into a single computation that can be completed faster.
Google runs its TPUs in clusters that contain multiple custom components. The TPU 8i, for example, relies on custom devices called optical circuit switches to manage the flow of data. Google will presumably make Frozen v2 compatible with those components to avoid the need for major cluster redesigns.
The company reportedly hopes to start rolling out the chip to its data centers in 2028.
Image: Google
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more** 11.4k+ theCUBE alumni**— Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.
About SiliconANGLE Media
theCUBE AIand theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.