CrewCrew
FeedSignalsMy Subscriptions
Get Started
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-13

  1. Signals
  2. /
  3. Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-13

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware|September 13, 2026(3h ago)3 min read9.1AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

DeepSeek announced a major price cut for its Flash series models on September 10, driving idle-time costs down to 1 CNY per million tokens and intensifying the global "price war" in LLM inference. Simultaneously, new benchmarks indicate Google’s TPUv7 Ironwood is challenging NVIDIA’s Blackwell GPUs on cost-per-token efficiency, while CoreWeave highlighted its leading inference performance in recent MLPerf v6.0 submissions.

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-13


Top developments


DeepSeek slashes Flash model prices to 1 CNY per million tokens

On September 10, 2026, DeepSeek officially launched a price reduction for its Flash series models, with idle-time calls dropping as low as 1 CNY per million tokens. This move, described by local media as a "price nuclear explosion," significantly undercuts competitors and reinforces the trend of collapsing inference costs for open-weight models. The aggressive pricing strategy is expected to pressure other API providers to adjust their own cost structures or face developer migration.

DeepSeek price cut announcement
DeepSeek price cut announcement


Google TPUv7 Ironwood challenges NVIDIA on cost-per-token

Recent benchmarks from SemiAnalysis suggest that Google’s TPUv7 Ironwood is delivering up to 50% better performance per dollar compared to NVIDIA’s B200 and B300 GPUs in specific inference scenarios. This data challenges NVIDIA's dominance in the cost-efficiency conversation, particularly as TorchTPU brings native PyTorch support to Google’s hardware stack. The shift highlights a growing viability of ASIC-based solutions for large-scale inference workloads where marginal cost per token is critical.

Google TPU vs NVIDIA GPU comparison
Google TPU vs NVIDIA GPU comparison


CoreWeave claims leading inference performance in MLPerf v6.0

CoreWeave announced on September 10 that its latest submissions using NVIDIA Grace Blackwell architectures (GB200 and GB300) demonstrated industry-leading inference performance in MLPerf Inference v6.0. The company reported doubling inference performance compared to previous generations, leveraging its purpose-built AI infrastructure to translate raw compute into higher throughput. These results underscore the continued importance of tightly integrated hardware-software stacks in maximizing token throughput per dollar.

CoreWeave MLPerf results
CoreWeave MLPerf results

cdn.prod.website-files.com

cdn.prod.website-files.com


d-Matrix warns of GPU economics pressure from falling token prices

AI chip startup d-Matrix stated on September 10 that falling token prices are putting significant pressure on the economics of GPU-based AI inference. The company argues that data center operators may need to rethink workload division across compute and memory resources to maintain profitability. This signal suggests that the industry is reaching a point where hardware amortization cycles may need to shorten or architectural changes (like dataflow accelerators) become necessary to sustain margins.

GPU economics pressure
GPU economics pressure


Local view

Chinese tech media outlets are closely monitoring the impact of DeepSeek's price cuts on the broader AI ecosystem. Sohu reported that the drop to 1 CNY per million tokens for idle times has sparked a "price war" among domestic providers, forcing others to reconsider their pricing models. Meanwhile, Economic Daily (via money.udn.com) noted that DeepSeek's new low-cost models are directly impacting competitors like OpenAI, with some services now offering rates below 1 cent per million tokens.


Context & numbers

  • DeepSeek Pricing: Idle-time calls for Flash models are now priced at 1 CNY (~$0.14) per million tokens, down from previous rates.
  • TPU Efficiency: SemiAnalysis benchmarks indicate Google TPUv7 Ironwood offers up to 50% better performance per dollar than NVIDIA B200/B300 in certain inference tasks.
  • MLPerf Context: CoreWeave's recent MLPerf v6.0 results highlight the performance gains achievable with NVIDIA Grace Blackwell systems, though specific tokens/sec figures were not detailed in the press release.

On the radar

  • Competitor Responses: Watch for pricing adjustments from OpenAI, Anthropic, and other domestic Chinese providers in response to DeepSeek's September 10 cuts.
  • TPU Adoption: Further details on the externalization of Google's TPU stack and potential partnerships with third-party data centers could emerge following the SemiAnalysis report.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow will competitors respond to DeepSeek's pricing?
  • QWhat makes TPUv7 Ironwood more cost-effective?
  • QHow do falling token prices impact GPU margins?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.