Google Designs a Gemini-Specific AI Chip Called Frozen v2

Google is designing Frozen v2, a chip that hardwires parts of its Gemini model into silicon and could serve six to ten times more tokens per watt, targeting deployment by 2028.

By Olivia Grant Edited by Maria Konash Published:
Google Designs a Gemini-Specific AI Chip Called Frozen v2
Google is designing Frozen v2, a chip that embeds parts of its Gemini model into silicon for up to 10 times better efficiency. Image: Rubaitul Azad / Unsplash

Google is developing a new server chip that would embed parts of its Gemini AI model directly into the silicon to run the model more efficiently, The Information reported on July 20, citing people familiar with the plans.

The chip, internally dubbed “Frozen v2,” could serve between six and ten times more tokens per unit of power than Google’s latest custom AI chips, according to the report. The idea is to hardwire portions of Gemini’s architecture into the hardware itself, cutting the number of calculations the chip must perform and the distance data must travel to generate a response.

Google did not confirm or deny the report, saying only that its teams constantly experiment with new innovations and that not every project reaches production.

The distinction that makes Frozen v2 notable is what “efficiency” means here. The gains are specifically in inference, the token-by-token work of generating answers for users, which has become one of the largest ongoing costs of running AI at scale, rather than in training.

Most AI accelerators, including Google’s own Tensor Processing Units, are general-purpose enough to run many different models. Frozen v2 takes the opposite approach, optimizing for one model architecture by “freezing” some of its structure into the chip, much like a purpose-built appliance outperforms a general computer at a single task. Engineers are reportedly still deciding how much of Gemini to bake in.

Crucially, the chip would complement Google’s TPUs rather than replace them, and Google reportedly views it partly as a trial run, not something to produce at TPU scale.

The project is a response to a real and pressing constraint. The report says Google is grappling with a severe internal compute shortage that has fueled tensions and forced Google Cloud to turn away some outside business, an acute enough squeeze that Google agreed last month to pay SpaceX nearly $1 billion a month for bridge capacity.

A large jump in tokens per watt would lower the marginal cost of serving Gemini and let Google handle more demand without matching increases in electricity and data center capacity, potentially improving the economics of its AI services and freeing it to take on customers it has had to decline. News of the chip helped lift Alphabet shares about 3% ahead of its earnings report this week.

The Vertical Integration Bet

Frozen v2 represents a deeper stage of the industry’s move toward custom silicon, going beyond building general AI accelerators to designing chips tailored to a specific foundation model. It reflects Google’s long-standing full-stack strategy of co-designing hardware and software, and it fits a broader scramble as AI companies build their own chips to cut costs and reduce dependence on Nvidia.

OpenAI unveiled its first custom inference chip, Jalapeño, built with Broadcom in June; Anthropic is reportedly in talks with Samsung on a chipmaking partnership; and Amazon, Microsoft and Meta all have in-house silicon efforts. As frontier models converge in capability, the ability to serve billions of answers within realistic power and cost limits is becoming as important a differentiator as raw model quality, shifting a key front of the AI race into semiconductor design.

The Trade-Off and the Caveats

The approach carries a clear cost: flexibility. By hardwiring Gemini’s architecture into the chip, Frozen v2 works well only if Google keeps building models on the same underlying design, so a major architectural shift could render the specialized hardware far less useful, which is likely why Google is treating it as a limited experiment.

Several caveats also temper the excitement. This is a report based on anonymous sources, not a Google announcement, and the six-to-tenfold efficiency figure is an internal engineering projection for a chip still being designed, with real-world results unproven.

The payoff is also years away, targeted for 2028, so it does nothing for today’s capacity crunch, and the claimed multiple must hold up against a moving baseline as every rival, including Nvidia, targets better performance per watt. The reported timing is notable given Google’s more immediate pressures, including a delayed next Gemini Pro release, senior researchers departing for rivals, and Chinese models now reportedly accounting for a large and growing share of US enterprise token use.

Frozen v2 is a credible long-term bet on efficiency, but for now it is a promising direction rather than a shipping solution.