
Reports Surface Google's Frozen v2 Chip
TechCrunch reported on 20 July 2026 that Google is developing a custom server chip internally codenamed Frozen v2, citing The Information, which originally broke the story. The chip is being designed to run Gemini AI inference workloads at six to ten times the efficiency of Google's current tensor processing units, measured in tokens generated per watt of power consumed. According to the reports, Google engineers working on the project are treating it as an exploratory effort, and the company has not announced plans to manufacture Frozen v2 at the scale of its production TPU fleet. The internal release target is 2028. Despite the early-stage framing, the efficiency projection — a potential tenfold improvement in tokens per watt — and the strategic context in which it sits make this worth understanding for teams thinking about the long-range trajectory of AI inference economics.
What Frozen v2 Actually Does
The core design philosophy of Frozen v2 is to hardwire portions of Gemini's architecture directly into silicon. General-purpose AI chips, including Google's own TPUs and Nvidia's production GPU lines, are flexible: they can run a wide variety of neural network architectures efficiently because their transistor layouts and memory hierarchies are designed for broad workload compatibility. A chip that hardwires a specific architecture gives up that flexibility in exchange for much higher efficiency on the one workload it was designed for. Frozen v2 would bake in the matrix factorisation patterns, attention configurations, and memory access sequences that Gemini uses during inference, eliminating a large fraction of the computation and data movement that a general-purpose chip must perform to achieve the same result. The tradeoff is that Frozen v2 would be of limited use for any model architecture other than Gemini.
The Name and the Design Bet
The "Frozen" name reflects a specific hypothesis: that the Gemini architecture has stabilised sufficiently to justify hardwiring it into silicon. Chip design and manufacturing take years. A company that invests in hardwiring an architecture into silicon is betting that the fundamental topology of the model will not change so dramatically between the chip's design completion and its production deployment that the efficiency gains are eroded by architectural drift. The Frozen v2 project is therefore implicitly a statement that Google's internal view is that the core Gemini architecture is settling — even as the company continues to iterate on training data, fine-tuning, and model scale.
The Compute Shortage Context
The proximate motivation for Frozen v2, according to the reporting, is an internal compute shortage at Google Cloud. Demand for Gemini inference through Google Cloud's enterprise offerings has outpaced the rate at which Google can supply TPU capacity to serve it. Enterprise customers wanting to run large-scale Gemini workloads have encountered capacity constraints that limit the amount of inference they can purchase. A chip that delivers six to ten times the inference per watt does not merely lower running costs — it allows Google to serve substantially more enterprise inference demand from the same physical data centre footprint and power envelope, without waiting for new data centre construction to complete.
This context is relevant for cloud infrastructure planning. If Google successfully deploys Frozen v2 at meaningful scale by 2028, the cost per Gemini inference token via Google Cloud would fall substantially, and the capacity constraints that have pushed some enterprise customers toward competitors could ease. For organisations currently making multi-year cloud AI infrastructure commitments, the potential change in Google Cloud's inference economics is a variable worth tracking.
The Broader Custom Silicon Race
Frozen v2 places Google in a race that other major AI labs are also running. Meta has been manufacturing its MTIA inference chip in volume since 2025. Amazon has deployed Trainium 2 and Inferentia 3 at scale on AWS for internal and customer workloads. Microsoft is understood to have inference-specific silicon projects at various stages. The general pattern is that as AI inference volumes grow large enough, the economics of custom silicon become compelling: a chip optimised for a specific model family can run that family substantially more cheaply than a general-purpose accelerator, and at the inference volumes these companies now run, even a 20 per cent efficiency improvement translates into hundreds of millions of dollars in annual savings. A six-to-tenfold gain, if realised, would be transformative.
What This Means for Indian Teams Building on Google Cloud
For Indian software companies and enterprise teams with multi-year cloud and AI infrastructure strategies, the Frozen v2 report matters for two reasons. First, it reinforces the direction that Google has chosen for Gemini's economic trajectory: if the 6-10x efficiency projection is even partially realised, inference costs on Vertex AI and Google Cloud would fall significantly, making high-volume Gemini integrations more commercially viable for cost-sensitive MSME-facing and consumer-scale applications where current per-token costs constrain deployment economics. Second, it is a reminder that AI infrastructure economics are a moving target. Teams evaluating multi-year commitments to a single cloud AI provider should account for the possibility that efficiency improvements arriving in 2027 and 2028 could materially shift the comparative economics between providers — and that locking in long-term contracts before those shifts occur carries its own risk.
The Bottom Line
Reports published on 20 July 2026 describe Google's internal Frozen v2 project: a custom server chip designed specifically for Gemini inference at six to ten times the efficiency of Google's current TPUs. The chip hardwires Gemini's architecture into silicon, trading architectural flexibility for dramatically lower computation and data movement per token. Google engineers are treating it as an exploratory project and the company has not committed to manufacturing it at production scale. The internal target date is 2028. The project is driven by a Google Cloud compute shortage limiting enterprise Gemini supply. If realised at scale, a Frozen v2 deployment would substantially reduce Gemini inference costs and ease capacity constraints that have challenged some Google Cloud enterprise customers.
Frequently Asked Questions
What is Google's Frozen v2 chip?+
Frozen v2 is an internal Google project to build a custom server chip optimised specifically for running Gemini AI inference workloads. TechCrunch reported on 20 July 2026, citing The Information, that Google engineers working on the project project it could deliver six to ten times the inference efficiency of Google's current tensor processing units, measured in tokens generated per watt. The chip achieves this by hardwiring portions of Gemini's architecture directly into silicon, eliminating computation and data movement that general-purpose chips must perform. Frozen v2 is internally treated as an exploratory project. Google has not committed to manufacturing it at the scale of its production TPU fleet. The internal target date is 2028.
How does hardwiring an AI architecture into silicon improve efficiency?+
General-purpose AI chips, including Nvidia GPUs and Google TPUs, are designed to execute a wide variety of neural network architectures efficiently. That flexibility requires circuitry and memory hierarchies that handle multiple possible computation patterns. When a chip is designed for a specific architecture — hardwiring the attention configurations, matrix factorisation patterns, and memory access sequences that the target model uses — it can eliminate the flexibility overhead that a general-purpose chip carries. The result is substantially more efficient inference: the same amount of silicon and power produces significantly more output tokens per second. The tradeoff is that a hardwired chip delivers little or no benefit for model architectures other than the one it was designed for.
Why is Google building Frozen v2 now?+
The proximate motivation for Frozen v2, according to TechCrunch citing The Information on 20 July 2026, is an internal compute shortage at Google Cloud. Demand for enterprise Gemini inference has been outpacing the rate at which Google can supply TPU capacity, leaving some enterprise customers unable to access the volume of Gemini inference they require. A chip delivering six to ten times the tokens per watt would allow Google to serve substantially more enterprise demand from the same data centre footprint and power budget, without waiting for new facility construction. The project also reflects a strategic bet that the Gemini architecture has stabilised sufficiently to justify hardwiring it into silicon — a multi-year commitment that would be wasted if the architecture changed substantially before production deployment.
What does the Frozen v2 chip mean for enterprise cloud customers using Gemini on Google Cloud?+
Frozen v2 is a 2028 internal project, so no near-term change to Google Cloud Gemini pricing or capacity is expected. However, if the project reaches production scale by 2028, it would have two material effects for enterprise customers. First, per-token inference costs via Vertex AI and Google Cloud's Gemini endpoints would fall substantially, making high-volume applications more commercially viable. Second, the compute shortage that has constrained enterprise Gemini supply would ease as Google's effective data centre inference capacity multiplies. For organisations planning multi-year AI infrastructure commitments, the potential shift in Google Cloud's AI inference economics by 2028 is a relevant planning variable, though it should be weighed against the uncertainty inherent in any multi-year hardware development roadmap.
Written by
TechPillow Team
Sharing insights on technology, product development, and the Indian tech ecosystem.