Google is building an AI chip that does something genuinely unusual: it takes pieces of its Gemini model and embeds them directly into the silicon itself. The project, internally called “Frozen v2,” represents a distinct branch from Google’s existing Tensor Processing Unit lineup and could deliver efficiency gains of six to ten times over current TPUs.
What Google is actually building
The core idea behind Frozen v2 is straightforward in concept, even if the engineering is anything but. Rather than running Gemini’s model weights and computations on general-purpose AI accelerator silicon, Google plans to hardcode specific elements of the model architecture directly into the chip’s design. Instead of the chip being a flexible calculator that can run any AI model, parts of the chip would be purpose-built to run Gemini and nothing else.
The projected 6-10x efficiency improvement over Google’s latest TPUs is a staggering number. Google’s latest generation consists of the TPU 8t for training and TPU 8i for inference, announced in 2026. Frozen v2 isn’t meant to replace those chips. It’s a parallel track, purpose-built for a specific job. Deployment is expected as early as 2028.
Google’s long history with custom silicon
The company has been developing custom silicon solutions since the mid-2010s, with its first TPU publicly revealed in 2018. The original motivation was practical: Google’s internal workloads for Search and YouTube demanded more efficient processing than off-the-shelf GPUs could deliver at reasonable cost and power consumption.
Since then, the TPU family has evolved considerably. Early generations focused primarily on inference. Subsequent iterations, including chips codenamed Ironwood and Trillium, expanded to handle both training and inference at scale and have been used for Gemini’s development and deployment across Google’s product ecosystem.
Frozen v2 is model-specific by design, whereas previous TPU generations were general-purpose AI accelerators flexible enough to run various models and workloads. Hardcoding model elements into a chip means those elements can’t easily be changed later.
The broader industry context
Amazon has its Trainium and Inferentia chips. Microsoft has Maia. Meta has been developing custom silicon for its AI workloads. But what Google is proposing with Frozen v2 goes a step further than what any competitor has publicly announced. Embedding model-specific logic into silicon is a more radical approach than simply designing a better general-purpose AI accelerator.
What this means for investors and markets
As of mid-2025, no blockchain integration, token angle, or decentralized compute play has emerged around Frozen v2, and the chip has not appeared in crypto-native media as of July 20, 2026.
Efficiency breakthroughs in AI inference hardware could lower the cost of running AI models at scale, with downstream effects on decentralized AI compute networks, on-chain AI agents, and AI-driven trading systems, all of which depend on the cost structure of running inference workloads. Google’s willingness to invest in model-specific silicon suggests the company views Gemini as a long-term franchise worth defending with substantial hardware R&D investment.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

19 hours ago
28









English (US) ·