Google TPU Ironwood Beats NVIDIA B200 by 50% in Cost-Per-Token Inference

Google TPU Ironwood leads NVIDIA B200 by 50% in inference performance per dollar, costing $0.181 per million tokens. See how it compares to B300.

Google TPU Ironwood
Google TPU Ironwood

Google released benchmark data for its seventh-generation TPU Ironwood, highlighting a significant cost advantage over ’s latest GPUs. This shift matters because AI inference workloads are increasingly driven by price-per-token metrics rather than raw peak performance alone. Buyers and cloud providers now have concrete evidence that specialized silicon can outperform general-purpose accelerators in efficiency.

Google releases benchmark data showing TPU Ironwood offers superior inference efficiency compared to NVIDIA GPUs

The TPU Ironwood utilizes a 256×256 systolic array matrix unit and supports native FP8 precision. Google tested this hardware using the Qwen3.5 397B FP8 model with 8000 input and 1000 output tokens. The TPU Ironwood features 192GB of HBM memory and a 256×256 systolic array matrix unit that natively supports FP8 operations.

  • Memory: 192GB HBM
  • Matrix Unit: 256×256 systolic array
  • Native Precision: FP8
  • Cost per Million Tokens (TPU): $0.181
  • Cost per Million Tokens (B200): $0.222

Performance comparisons show TPU Ironwood leads NVIDIA B200 by up to 50% in inference performance per dollar. The chip also leads NVIDIA B300 by 96% in the same metric. Physical throughput only leads by approximately 5%, but the lower rental cost amplifies the per-dollar advantage significantly.

Cost analysis reveals TPU Ironwood charges $0.181 per million tokens. This is lower than NVIDIA B200 at $0.222 and NVIDIA B300 at $0.276. Google also introduced TorchTPU to help PyTorch users run on TPU with minimal code changes. Anthropic has already purchased over 1 million TPUs and is expected to become the largest single TPU user by 2029.

We looked at Google Gemini 3.5 Flash Adds Built-In Computer Use in our earlier Google coverage. The new benchmark data reinforces Google's strategy of competing on total cost of ownership rather than just raw speed. This approach directly impacts how enterprises evaluate AI infrastructure spending.

Google continues to expand its TPU ecosystem with software tools and competitive pricing. The data confirms that TPU Ironwood offers a viable alternative to NVIDIA GPUs for cost-sensitive inference tasks. Companies prioritizing token efficiency should consider this architecture for their deployment pipelines.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion