In brief
- Google is reportedly developing a server chip called Frozen v2 that hardwires part of Gemini's architecture into silicon.
- Based on reports, engineers project six to ten times the efficiency of current TPUs and a 2028 deployment target.
- Alphabet shares climbed roughly 3% during Monday's trading session after the news broke, ahead of Q2 2026 earnings due Wednesday, July 22.
Google is building a chip designed for one job: running Gemini faster and cheaper.
The chip, codenamed Frozen v2, was reported by The Information on Monday and gives Google a potential answer to a problem it can't spend its way out of fast enough: It is running out of capacity to serve the AI demand it has already generated.
In March, Google told Meta it couldn't fill the volume of Gemini compute Meta wanted to purchase. Meta had to instruct employees to ration their AI usage. Google—spending up to $190 billion on AI infrastructure this year—was turning away customers because it didn't have enough servers to serve them.
So now it's building a chip designed only for its AI models.
There’s not much information about this new chip, but by naming convention it’s not another upgrade to Google's Tensor Processing Units (TPUs)—the custom chips Google has been building since 2015 that power Gemini and its Cloud services for outside developers.
Those Tensor chips run any AI model loaded onto them. Frozen v2 does something different. Per the reports, it bakes part of Gemini's architecture—the structural blueprint that determines how the model routes and processes information—directly into the hardware.
In machine learning, "freezing" means locking something permanently in place. Here, what gets frozen is the architecture, not the model's weights (the actual knowledge Gemini picks up through training, which stays updatable). By hardwiring this blueprint into the chip's circuits, the chip skips redundant calculations and stops shuttling data across memory on every query. Engineers project a six to ten times improvement in tokens—the small text chunks that make up each AI response—generated per watt of electricity consumed.
That's the difference between Google serving ten queries for the power cost of one.
If you use Gemini, Frozen v2 won't change how it feels to you. But it changes what it costs to run—and a cheaper-to-run Gemini competes harder against OpenAI, Anthropic, and Chinese labs that already account for up to 45% of U.S. company AI token usage, largely because they run 60–90% cheaper. You may not have cheaper AI, but Google will likely be more profitable. Alphabet shares climbed roughly 3% during Monday's session on the news, touching $356 intraday. The company reports Q2 2026 earnings on Wednesday, July 22, and the pump receded in today’s session as investors wait for Google’s most recent results..
This is yet another effort by a major AI company to kill its over-reliance on Nvidia hardware to develop its products. Nvidia controls roughly 85% of the GPU market for AI, and every major tech company wants out.
Nvidia's hardware was originally built for video games, not language models—it works, just with overhead that purpose-built chips don't carry. At Google's scale, a 6–10x efficiency gap isn't abstract. It's billions of dollars. Meta, Amazon, Microsoft, and OpenAI all have custom silicon programs for exactly that reason.
As Decrypt reported in March, even AWS—which committed to deploying 1 million Nvidia GPUs through 2027—is building its own chips simultaneously to cut that long-term exposure.
Frozen v2 is still exploratory. Key design decisions aren't finalized, Google hasn't confirmed the project exists, and the chip won't be offered to outside Cloud customers—hardware hardwired for one model can't run anyone else's. Deployment is targeted for 2028 at the earliest, according to reports.
In the meantime, Google is paying SpaceX $920 million a month to rent 110,000 Nvidia GPUs from xAI's data centers as a bridge.