Google's 'Frozen v2' chip would bake Gemini directly into silicon for a 6–10x efficiency leap
Alphabet is reportedly designing a server chip that hardcodes parts of Gemini's architecture into hardware, targeting deployment as early as 2028.
What matters
- Google is developing a server chip internally called 'Frozen v2' that hardcodes parts of Gemini's architecture directly into silicon.
- Engineers estimate the chip could deliver 6–10x more tokens per watt than Google's latest TPUs.
- The chip is a specialized branch separate from Google's TPU lineup, not a replacement.
- Google is targeting deployment as early as 2028, but the design is still being finalized.
- The project was partly motivated by a severe internal AI compute capacity crunch that led Google Cloud to decline external customer business.
What happened
Google is developing a new server chip, internally referred to as "Frozen v2," that would embed parts of its Gemini AI model architecture directly into the silicon itself, according to a report by The Information picked up by CNBC, Reuters, and Bloomberg. Alphabet shares rose as much as 3.7% on the news.
Unlike Google's existing Tensor Processing Units (TPUs), which are general-purpose accelerators designed to run a wide range of AI workloads, Frozen v2 takes a specialized approach. It would hardcode specific elements of Gemini's neural-network architecture into the chip's circuitry, cutting down the number of calculations the hardware must perform and how far data must travel to generate a response. Engineers can still refresh the model by loading new weights, but the underlying structure stays fixed—or "frozen."
Engineers working on the project estimate the chip could deliver six to ten times the token output per watt compared with Google's latest TPUs. Google is targeting deployment as early as 2028, though engineers are still finalizing the chip's design and deciding how much model information will be hardwired into it. Google has not confirmed the project.
Why it matters
The reported project is a serious bet on where AI infrastructure goes next. Today's AI chips keep the model in memory and shuttle its data back and forth—a flexible approach that costs power and time. Frozen v2 would flip that model: instead of the chip being a flexible calculator that can run any AI model, parts of the chip would be purpose-built to run Gemini and nothing else.
The efficiency stakes are enormous. A 6–10x improvement in tokens served per unit of power would dramatically lower the cost of running Gemini at scale. Google reportedly developed the project in part to relieve a severe internal capacity crunch that has created friction within the company and led Google Cloud to decline business from external customers.
The trade-off is flexibility. Because parts of the model are hardwired into the silicon, the chip can only support future Gemini versions if Google keeps its foundational architecture intact. That means Frozen v2 is not a replacement for Google's TPUs—it's a parallel track, purpose-built for a specific job.
What to watch
- Deployment timeline: Google is targeting 2028, but the chip is still in the design phase. How much of Gemini gets hardwired—and whether that architecture remains stable enough across future model versions—will determine whether the bet pays off.
- Capacity crunch signals: If Google Cloud continues turning away external customers due to compute shortages, that adds urgency to the project and could signal broader AI infrastructure bottlenecks across the industry.
- Competitive ripple effects: A chip that bakes a model into silicon challenges the assumption that general-purpose accelerators like Nvidia's GPUs will dominate AI inference. Watch for whether other labs pursue similar model-specific hardware strategies.
- Google's confirmation: The company has not officially confirmed the project. Any public acknowledgment—or silence—will be telling.
What to do next
Developers
Monitor Google Cloud's TPU and Gemini API roadmap for any signals about model-specific hardware tiers or new inference pricing tied to efficiency gains.
If Frozen v2 deploys, it could eventually lower Gemini inference costs and unlock new serving tiers—but it won't be available until 2028 at the earliest.
Founders
Assess whether model-specific hardware trends could affect your cloud AI spend and vendor lock-in strategy over the next 3–5 years.
If hyperscalers begin offering model-baked silicon at lower cost per token, it could shift the economics of running AI workloads and influence which provider you build on.
PMs
Track Google's internal capacity crunch as a leading indicator of broader AI compute scarcity and plan contingency providers for latency-sensitive AI features.
Google Cloud reportedly declining external business due to compute shortages signals that capacity constraints are real and could affect service availability.
Investors
Weigh the strategic implications of model-specific silicon for the GPU-dominated AI hardware thesis, and watch for similar bets from other labs.
A 6–10x efficiency claim, if realized, would challenge assumptions about general-purpose accelerator dominance and could reshape the competitive landscape for AI infrastructure stocks.
Operators
Review your AI inference cost-per-token benchmarks and model whether a 6–10x efficiency improvement would materially change your unit economics.
Understanding the potential impact of model-specific hardware on serving costs helps you prepare for a shift in infrastructure economics, even if deployment is years away.
Testing notes
Caveats
- Frozen v2 is an unconfirmed internal project with no public hardware, SDK, or API available. Deployment is targeted for 2028 at the earliest, and the chip design is still being finalized. There is nothing to test at this time.