ProductApr 23, 2026

Google Cloud splits its TPU line with two new chips for the agentic era

What’s the deal? Google CloudDealroom has a profile for this one. Try Dealroom → has unveiled its eighth-generation Tensor Processing Units in two distinct versions — a first in the chip’s decade-long history. The TPU 8t is built for training frontier AI models; the TPU 8i is purpose-built for inference. Both were announced at Google Cloud Next in Las Vegas on April 22, 2026, alongside a $750M fund to accelerate corporate AI adoption.

The split reflects a fundamental shift in AI computing. Inference — running models after they are trained — now generates the bulk of demand, driven by the rise of AI agents that trigger 20 to 50 times more compute transactions than a standard chatbot query.

Why now? The AI chip race has become an inference race. Nvidia in March 2026 debuted its own inference solution pairing GPUs with licensed technology from Groq, acquired for $20B. Cerebras Systems struck a major deal with Amazon Web ServicesDealroom has a profile for this one. Try Dealroom → and filed for an IPO targeting May 2026. Microsoft and AmazonDealroom has a profile for this one. Try Dealroom → are both building custom inference silicon.

Google Cloud chief executive Thomas KurianDealroom has a profile for this one. Try Dealroom → put it plainly: “If you don’t have inference, you cannot cover the cost of your training. Eventually inference is going to be at least as big, if not bigger, than the training market.”

The TPU 8t delivers nearly 3x the compute per pod over the previous generation (Ironwood), scaling to 9,600 chips and two petabytes of shared memory. The TPU 8i tackles the “memory wall” with 288GB of high-bandwidth memory and 3x more on-chip SRAM, cutting chip-to-chip latency by over 50% via a new Boardfly topology. Both chips offer 2x better performance-per-watt than their predecessors.

What could go wrong? Google is not replacing Nvidia — it will still offer Nvidia’s Vera Rubin GPUs through its cloud. Despite a decade of TPU development, Nvidia remains a nearly $5T company. Analyst Patrick Moorhead noted he predicted TPUs would threaten Nvidia back in 2016; that bet has not paid off.

Porting workloads from GPUs to TPUs still requires effort, though Google has expanded framework support to include PyTorch, vLLM, and SGLang.

The signal: The era of the general-purpose AI chip is fading. Google’s decision to split training and inference into separate architectures signals that workload specialisation is now the path to meaningful efficiency gains — not incremental improvements to a single design.

Major external customers already include AnthropicDealroom has a profile for this one. Try Dealroom → (roughly one million chips), Meta Platforms, and Citadel Securities, which reported a 30% cost reduction. Morgan Stanley estimates 500,000 TPU sales could add roughly $13B in revenue by 2027. Google also promoted Amin VahdatDealroom has a profile for this one. Try Dealroom → to chief technologist of AI infrastructure, reporting directly to chief executive Sundar PichaiDealroom has a profile for this one. Try Dealroom → — a sign that custom silicon is now a C-suite priority, not just an engineering project.

Sources:
Google Blog
Wall Street Journal
Bloomberg
TechCrunch
Forbes
Business Insider
Morningstar

Image credit:
Google

J,V.

Source: dealroom

More top stories