TL;DR
The News: Google is reportedly developing**"Frozen v2,"** a dedicated inference chip built specifically to run Gemini models.The Promise: Claims a6x to 10x increase in energy efficiency by hardcoding core Gemini operational logic directly onto silicon.The Catch: Highly specialized ASICs trade flexibility for speed—major changes to Gemini’s architecture could make the chip obsolete.
Large artificial intelligence models require significant computing power, especially when generating a response. This phase, called inference, uses specialized processors in data centers. Its cost depends notably on the number of queries, their length, and the efficiency of the hardware.
The project reportedly involves integrating a fixed part of Gemini's operation into the silicon. Unlike a very versatile chip, such a specialized accelerator can avoid certain calculations or data transfers. This would significantly reduce the energy needed to generate each unit of text.
This promise, however, must remain conditional. The published information relies on anonymous sources, and Google has not detailed the architecture, performance, or release timeline. The estimates mentioned in the press are therefore not results verified by the company.
Google already has its own TPUs, chips designed for machine learning. Frozen v2 would fit into this strategy of hardware integration, as computing needs for generative AI increase rapidly. The stakes are technical, but also economic.
A component optimized for a specific model can be very efficient, but less flexible than a general-purpose processor. Any major evolution of Gemini could require new hardware design, and therefore the production of new chips. This is the usual trade-off between specialized performance and adaptability.
If this project comes to fruition, its value will be measured especially at the data center scale. Lower energy consumption per response could reduce operating costs and free up computing capacity for more powerful models.