Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-12
Google’s TPUv7 Ironwood is challenging NVIDIA’s dominance in inference cost-efficiency, with SemiAnalysis reporting up to 50% better performance per dollar compared to B200/B300 GPUs. Meanwhile, DeepSeek executed a significant price cut for its Flash models on September 10, driving cached token costs down to $0.02 per million tokens and intensifying the global cost-per-token war.
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-12
Top developments
Google TPUv7 Ironwood Outperforms NVIDIA on Cost-Per-Token
SemiAnalysis released new benchmarks indicating that Google’s TPUv7 (Ironwood) delivers up to 50% better performance per dollar than NVIDIA’s B200 and B300 GPUs for inference workloads. This shift is accelerated by the rapid externalization of the TPU stack and the introduction of TorchTPU, which provides native PyTorch support, reducing the software barrier for adopting TPUs. The move threatens to erode NVIDIA’s CUDA moat by offering a compelling alternative for cost-sensitive AI operators.

DeepSeek Flash Models Slash Prices to Near-Zero Levels
On September 10, 2026, DeepSeek implemented a massive price reduction for its Flash series models, effective immediately. The new pricing structure sees idle-time cached input tokens drop to ¥0.02 (approx. $0.0028) per million tokens, while peak-time rates remain higher but competitive. This "price nuclear explosion," as described by local media, aims to capture developer mindshare and undercut competitors like OpenAI and Anthropic, particularly for high-volume, latency-tolerant workloads.
d-Matrix Warns of GPU Economics Pressure from Falling Token Prices
AI chip startup d-Matrix highlighted that the relentless decline in token prices is fundamentally altering the economics of GPU-based inference. As per-token costs fall, data center operators are forced to rethink workload distribution between compute and memory resources to maintain profitability. This trend favors specialized architectures that can deliver higher throughput at lower power costs, potentially accelerating the shift away from general-purpose GPUs for specific inference tasks.
Independent Benchmarks Compare GLM 5.3 Flash and DeepSeek V4 Flash
Yotta Labs published a direct comparison of GLM 5.3 Flash and DeepSeek V4 Flash, noting that while both are open MoE (Mixture of Experts) models, they have vastly different hardware requirements and cost structures. The analysis highlights that hardware bills vary significantly despite similar benchmark scores, emphasizing that "tokens per dollar" is heavily dependent on the underlying serving stack and hardware optimization rather than just model architecture.
Local view
Chinese tech media outlets such as Sohu and CSDN are closely tracking DeepSeek's aggressive pricing strategy. Sohu described the September 10 price cut as a "price nuclear explosion" that has sent shockwaves through the AI developer community, with many viewing it as a strategic move to dominate the low-cost inference market. CSDN analysts note that while unit prices have dropped, total bills for some users may rise due to increased usage volume, prompting a shift in how developers budget for AI services.
Context & numbers
- DeepSeek Pricing: Idle-time cached input is now ¥0.02/million tokens; peak-time input is ¥0.10/million tokens.
- TPU Efficiency: SemiAnalysis claims TPUv7 offers up to 50% better performance per dollar vs. NVIDIA B200/B300.
- Market Trend: Global LLM inference average prices have fallen significantly in 2026, with some analyses citing an 88% drop in average unit prices compared to previous years, though total spend continues to rise due to volume.
On the radar
- TPU Externalization: Watch for further announcements from Google regarding the availability of TPUv8i and expanded third-party access to the TPU stack, which could further commoditize inference hardware.
- Competitor Responses: Expect OpenAI and Anthropic to respond to DeepSeek's price cuts, potentially adjusting their own pricing tiers or introducing new efficiency-focused models.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.