Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-17
MLCommons released MLPerf Inference v6.1 results on September 16, marking the first peer-reviewed performance data for NVIDIA’s Vera Rubin NVL72 and introducing new Agentic and End-to-End RAG benchmarks. The round saw record participation with 30 submitters and 120 systems, highlighting significant software-driven efficiency gains on existing Blackwell hardware alongside debut hardware metrics. Meanwhile, Chinese market dynamics show DeepSeek aggressively cutting token prices to ~2 cents per million tokens for cached idle requests, intensifying cost-per-token competition.
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-17
Top developments
MLPerf v6.1 Debuts Vera Rubin and New Agentic Benchmarks
On September 16, 2026, MLCommons published MLPerf Inference v6.1 results, introducing two new tests: Agentic Inference and End-to-End RAG. This round set a participation record with 30 submitters and 120 systems, analyzing where the industry is investing in inference efficiency. The results provide the first peer-reviewed numbers for NVIDIA’s Vera Rubin NVL72, which achieved up to 3.7x better throughput than the previous GB300 NVL72 in specific configurations.

Software Optimization Yields Double-Digit Gains on Blackwell
Independent cloud providers reported 8 to 36 percent throughput gains on deployed NVIDIA Blackwell hardware through software optimization alone, without new hardware additions. This underscores that software stack maturity remains a critical lever for tokens-per-dollar efficiency, even as new ASICs and GPUs enter the market. Lambda.ai also reported pioneering agent workload results on datacenter hardware, including faster GPT-OSS 120B throughput.

AMD ROCm Optimizations Power MI355X Submissions
AMD submitted results using its ROCm stack, detailing optimizations for models like dlrm-v3, llama2-70b, and gpt-oss-120b. StorageReview noted a notable 512-GPU MI355X run in this round, contributing to the 5.7x per-accelerator gains observed across the benchmark suite. These submissions highlight the growing competitive pressure from AMD accelerators in the high-throughput inference segment.

Edge Agentic Benchmark Sees 6.4x Speedup on Jetson Thor
NVIDIA demonstrated a 6.4x speedup on the new MLPerf Edge Agentic benchmark using TensorRT Edge-LLM running on Jetson AGX Thor. This result is significant for moving AI agents from cloud data centers to edge devices like vehicles and robots, where latency and power efficiency are paramount over raw throughput.

Local view
In China, price competition has intensified with DeepSeek’s recent price cuts effective September 10, 2026. The Flash series model's idle-time price for cache hits dropped to approximately 2 cents per million tokens, a move described by local media as a "price nuclear explosion". Platforms like Alibaba Cloud have adopted dynamic pricing, though some users report their rates remain higher than official rates due to service tiers. Analysts note this shifts the competitive focus from pure model capability to infrastructure efficiency and token production costs, with DeepSeek planning a 1GW data center in Ulanqab to support this strategy.
Context & numbers
- Participation Records: MLPerf v6.1 featured 30 submitters and 120 systems, the broadest round to date.
- Efficiency Gains: Vera Rubin NVL72 showed up to 3.7x throughput improvement over GB300 NVL72 in peer-reviewed tests.
- Software Gains: Existing Blackwell hardware saw 8–36% throughput improvements via software updates alone.
- Token Pricing: DeepSeek Flash idle-time cache-hit pricing is now ~$0.02 per million tokens.
On the radar
- TPU Externalization: Reports indicate Google is accelerating the externalization of its TPU stack, with claims of up to 50% better performance per dollar compared to traditional GPUs in certain workloads.
- NVIDIA DSX MaxLPS: NVIDIA is promoting its DSX MaxLPS technology, claiming 40% more token throughput per megawatt on Vera Rubin platforms, aiming to address energy efficiency concerns in large-scale inference clusters.
- Speculative Decoding Adoption: Serving engines like vLLM and SGLang continue to integrate speculative decoding features, which can offer 2–5x latency improvements at the serving layer, becoming a standard expectation for efficient inference stacks.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.