Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-10
DeepSeek announced a massive price cut for its Flash model series, dropping idle-time costs to as low as 1 RMB per million tokens, signaling a new phase in the inference cost war. Meanwhile, independent benchmarks from Openbenchmarks highlighted Nebius as the leader in streaming throughput for GLM 5.3 Flash, while SemiAnalysis reported that Google's TPU externalization strategy is delivering up to 50% better performance per dollar compared to traditional GPU stacks.
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-10
Top developments

DeepSeek Flash Price Cut Hits 1 RMB per Million Tokens
On September 10, 2026, DeepSeek officially launched a full-scale price reduction for its Flash series models. The new pricing structure sees idle-time calls for the Flash model drop to 1 RMB per million tokens, a move described by local media as a "price nuclear explosion" that significantly lowers the barrier for high-volume AI application development. This aggressive pricing strategy forces competitors to re-evaluate their cost structures and may accelerate the adoption of Chinese open-weight models in global inference workloads where cost-per-token is the primary constraint.

Nebius Leads Independent Streaming Inference Benchmarks
In a benchmark conducted on September 3, 2026, Openbenchmarks evaluated ten LLM inference providers serving GLM 5.3 Flash. Nebius emerged as the leader with a median throughput of 269.3 tokens per second. The study also noted a p99 floor of 39.9 tokens per second and a time-to-first-token (TTFT) p95 of 7.11 seconds, providing concrete data on the latency-throughput trade-offs currently available in the market. These figures serve as a critical reality check against vendor claims, highlighting the variance in performance across different infrastructure providers.
Google TPU Externalization Claims 50% Better Performance per Dollar
SemiAnalysis reported on September 8, 2026, that the externalization of Google's TPU stack is gaining significant traction, with claims of up to 50% better performance per dollar compared to NVIDIA-based solutions. The report highlights the rapid deployment of Ironwood and TPUv8i chips, suggesting a strategic move to reduce the CUDA moat by offering high-efficiency alternatives for inference-heavy workloads. This development intensifies the hardware competition, particularly for enterprises looking to optimize long-term inference costs without being locked into a single vendor ecosystem.
Local view
Sohu and SMZDM are closely tracking the impact of DeepSeek's price cuts on the broader AI economy. Sohu characterized the event as a definitive shift in the "AI cost waterfall," noting that the new 1 RMB/MT idle rate is not just a promotional tactic but a structural change in how inference services are priced. Meanwhile, SMZDM published an analysis titled "Average Price Drops 88%, Bills Increase 10x," arguing that while unit prices have plummeted, the total consumption volume of LLMs has surged so dramatically that the overall market value continues to grow, requiring a complete recalibration of cost models for developers.
Context & numbers
- DeepSeek Flash Idle Price: 1 RMB per million tokens (effective Sept 10, 2026)
- Nebius Median Throughput: 269.3 tokens/sec for GLM 5.3 Flash (Sept 3, 2026)
- Nebius P99 Floor: 39.9 tokens/sec
- TPU Performance Claim: Up to 50% better performance per dollar vs. competitors
On the radar
- MLPerf Storage v3.0 Adoption: Following the release of MLPerf Storage v3.0 results in early September, expect more vendors to submit checkpointing and KV-cache optimized storage solutions, which are becoming critical bottlenecks for large-scale inference clusters.
- Speculative Decoding Standardization: With tools like vLLM and SGLang increasingly supporting speculative decoding, watch for new independent benchmarks specifically isolating the throughput gains from these techniques versus raw hardware improvements.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.