Google is reportedly developing its own AI chip codenamed Frozen v2, which firmly embeds the architecture of the language model Gemini into the circuitry, according to a report from the trade publication The Information. The chip is expected to deliver six to ten times more performance per watt than Google’s current TPU generation. The rollout in its own data centers is planned for no earlier than 2028.
Chip embeds Gemini architecture in hardware
Unlike TPUs, Google’s previous AI accelerators, Frozen v2 does not set the trained weights of the model in hardware, but rather its fundamental structure. The architecture thus remains fixed, while new weights can still be loaded. According to the report, this specification reduces the necessary computation steps per request and lowers the data traffic between memory and processing units – the six to ten times efficiency compared to current TPUs is independently unverified.
The chip is also expected to carry enough memory on board to run Gemini completely without accessing external RAM – this saves costly data transfers between components. Tom’s Hardware classifies this as a consistent continuation of Google’s TPU strategy: The current generation already separates training chips (TPU 8t) from inference chips (TPUi). Frozen v2 is expected to remain compatible with existing cluster technology, including optical switching systems between processing units. Google itself has not publicly commented on the plans so far.
Google aims for cheaper AI responses than competitors
The chip is intended solely for internal use. Unlike Google’s TPUs, which are also rented to external cloud customers like Meta, Frozen v2 is not intended to be sold. The report does not comment on costs or deployment in European data centers. At its core, it is about the cost per answered request: The cheaper a provider can offer inference – that is, the ongoing use of a fully trained model – the greater its margin will be.
If the promised efficiency leap is achieved, Google could offer Gemini responses at a lower cost than OpenAI and Anthropic, thereby gaining market share from them. Investors have already reacted to the report: Alphabet’s stock rose by 1.5 percent afterward. For companies integrating Gemini through Google Cloud or into their own software, such an efficiency leap could mean lower usage costs and faster response times in the medium term. This initiative comes at a time when Google, OpenAI, and Anthropic are increasingly competing over the cost per response rather than just model quality.
It remains to be seen whether Google will confirm the plans itself before the first rollout in 2028 – so far, all information comes from a single, unconfirmed report. It will also be crucial whether the promised efficiency actually materializes in the construction of real chips, as previous hardware announcements in the industry have often fallen short of their lab values. For Nvidia, whose graphics chips currently account for the majority of the AI training and inference market, a specialized Google chip would signal that competition is shifting more towards model-specific hardware.


