custom ai chip development

Splitting its AI silicon into dedicated training and inference paths, Google is reshaping its infrastructure around new custom Tensor Processing Units that position TPU 8t and TPU 8i as the backbone of an agentic-era AI platform. This bifurcated eighth-generation design marks a shift away from general-purpose accelerators toward specialized chips tuned for distinct phases of AI workloads. TPU 8t and 8i sit alongside Trillium, Google’s sixth-generation TPU, and Ironwood, its seventh-generation architecture, forming a family of custom processors tailored for large-scale Gemini systems. At Google Cloud Next 2026, TPU 8t and TPU 8i were unveiled as distinct chips for training and inference, highlighting their specialized TPU roles.

Co-developed with Google DeepMind and hosted on Axion ARM-based CPUs, these TPUs are assembled into tightly integrated AI supercomputing platforms intended to sustain frontier-scale models over multiple generations. The integration of AI-driven resource management enhances overall performance and efficiency in processing.

On the training side, Trillium extends Google’s matrix-multiply units and integrates third-generation SparseCore accelerators to better handle ultra-large embeddings and recommendation workloads, while delivering roughly 4.7x peak compute per chip compared with TPU v5e. TPU 8t builds on that foundation by scaling a single superpod to 9,600 chips and up to 2 petabytes of shared high-bandwidth memory, enabling synchronized training of trillion-parameter Gemini models in one logical system.

Each TPU 8t superpod can provide about 121 exaflops of native FP4 compute, yielding around three times the peak performance per superpod over the previous generation and compressing frontier model development cycles from months to weeks. Google targets up to 2.8x better training price-performance versus Ironwood, combining higher throughput with more efficient key-value cache handling to deliver approximately 2.5x better cost efficiency for Gemini 3 training.

TPU 8i is engineered specifically for post-training inference, high-concurrency reasoning, and reinforcement learning workloads characteristic of agentic AI systems that must respond quickly at scale. The chip triples on-chip SRAM to around 384 MB and doubles inter-chip interconnect bandwidth to roughly 19.2 Tb/s, directly improving latency-sensitive Gemini serving and long-context reasoning performance.

A new Boardfly network topology reduces network diameter by about 56 percent for mixture-of-experts and complex reasoning workloads, tightening end-to-end response times when models route tokens across thousands of experts. Together with the broader Trillium infrastructure, TPU 8i delivers around three times the inference throughput for generative models such as Gemma 2 and modern diffusion systems, and is described as achieving up to a tenfold performance gain versus seventh-generation configurations for large-scale AI queries.

Ironwood, the seventh-generation TPU, introduced native FP8 support and was architected for high-performance inference while retaining state-of-the-art training capability, establishing a bridge between dual-purpose accelerators and today’s specialized silicon. Ironwood pods can connect up to 9,216 chips through a proprietary Inter-Chip Interconnect operating near 9.6 Tb/s, forming a single logical AI supercomputer that set earlier benchmarks for Gemini-scale workloads.

Those capabilities now underpin the economics and topology lessons embedded in TPU 8t and 8i, which double interconnect bandwidth and substantially expand memory while focusing each chip on either training or inference.

Gemini performance accelerates overall.

You May Also Like

SK Group Warns Global AI Memory Shortage Is Becoming an Economic Security Risk

The global AI memory shortage is spiraling into an economic security crisis, and the consequences may be far worse than anyone anticipated.

Agentic AI Infrastructure Creates New Demand for High-Speed Memory and Storage

Beyond traditional compute needs, agentic AI is rewriting the rules of memory and storage infrastructure—and the implications are only beginning.

Nokia and NVIDIA Launch an AI-RAN Platform for the Future of Mobile Networks

The Nokia and NVIDIA AI-RAN platform promises revolutionary spectral efficiency gains, but the full scope of its impact on 6G remains to be seen.

AI Inference Chips Attract $400 Million as Investors Move Beyond GPU Financing

Capital is flooding into AI inference chips, with $400 million deals signaling a seismic shift away from GPUs—but who will dominate?