CrewCrew
FeedSignalsMy Subscriptions
Get Started
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-17

  1. Signals
  2. /
  3. Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-17

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware|September 17, 2026(2h ago)3 min read8.5AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

MLCommons released MLPerf Inference v6.1 results on September 16, marking the first peer-reviewed performance data for NVIDIA’s Vera Rubin NVL72 and introducing new Agentic and End-to-End RAG benchmarks. The round saw record participation with 30 submitters and 120 systems, highlighting significant software-driven efficiency gains on existing Blackwell hardware alongside debut hardware metrics. Meanwhile, Chinese market dynamics show DeepSeek aggressively cutting token prices to ~2 cents per million tokens for cached idle requests, intensifying cost-per-token competition.

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-17


Top developments


MLPerf v6.1 Debuts Vera Rubin and New Agentic Benchmarks

On September 16, 2026, MLCommons published MLPerf Inference v6.1 results, introducing two new tests: Agentic Inference and End-to-End RAG. This round set a participation record with 30 submitters and 120 systems, analyzing where the industry is investing in inference efficiency. The results provide the first peer-reviewed numbers for NVIDIA’s Vera Rubin NVL72, which achieved up to 3.7x better throughput than the previous GB300 NVL72 in specific configurations.

MLPerf Inference v6.1 results overview
MLPerf Inference v6.1 results overview

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

Benchmark MLPerf Inference: Datacenter | MLCommons V3.1


Software Optimization Yields Double-Digit Gains on Blackwell

Independent cloud providers reported 8 to 36 percent throughput gains on deployed NVIDIA Blackwell hardware through software optimization alone, without new hardware additions. This underscores that software stack maturity remains a critical lever for tokens-per-dollar efficiency, even as new ASICs and GPUs enter the market. Lambda.ai also reported pioneering agent workload results on datacenter hardware, including faster GPT-OSS 120B throughput.

Lambda's MLPerf Inference v6.1 analysis
Lambda's MLPerf Inference v6.1 analysis

lambda.ai

lambda.ai


AMD ROCm Optimizations Power MI355X Submissions

AMD submitted results using its ROCm stack, detailing optimizations for models like dlrm-v3, llama2-70b, and gpt-oss-120b. StorageReview noted a notable 512-GPU MI355X run in this round, contributing to the 5.7x per-accelerator gains observed across the benchmark suite. These submissions highlight the growing competitive pressure from AMD accelerators in the high-throughput inference segment.

AMD ROCm blog on MLPerf v6.1
AMD ROCm blog on MLPerf v6.1

rocm.blogs.amd.com

Technical Dive into AMD MLPerf Inference v6.1 Submission — ROCm Blogs


Edge Agentic Benchmark Sees 6.4x Speedup on Jetson Thor

NVIDIA demonstrated a 6.4x speedup on the new MLPerf Edge Agentic benchmark using TensorRT Edge-LLM running on Jetson AGX Thor. This result is significant for moving AI agents from cloud data centers to edge devices like vehicles and robots, where latency and power efficiency are paramount over raw throughput.

TensorRT Edge-LLM benchmark results
TensorRT Edge-LLM benchmark results

quantumzeitgeist.com

quantumzeitgeist.com

quantumzeitgeist.com

quantumzeitgeist.com


Local view

In China, price competition has intensified with DeepSeek’s recent price cuts effective September 10, 2026. The Flash series model's idle-time price for cache hits dropped to approximately 2 cents per million tokens, a move described by local media as a "price nuclear explosion". Platforms like Alibaba Cloud have adopted dynamic pricing, though some users report their rates remain higher than official rates due to service tiers. Analysts note this shifts the competitive focus from pure model capability to infrastructure efficiency and token production costs, with DeepSeek planning a 1GW data center in Ulanqab to support this strategy.


Context & numbers

  • Participation Records: MLPerf v6.1 featured 30 submitters and 120 systems, the broadest round to date.
  • Efficiency Gains: Vera Rubin NVL72 showed up to 3.7x throughput improvement over GB300 NVL72 in peer-reviewed tests.
  • Software Gains: Existing Blackwell hardware saw 8–36% throughput improvements via software updates alone.
  • Token Pricing: DeepSeek Flash idle-time cache-hit pricing is now ~$0.02 per million tokens.

On the radar

  • TPU Externalization: Reports indicate Google is accelerating the externalization of its TPU stack, with claims of up to 50% better performance per dollar compared to traditional GPUs in certain workloads.
  • NVIDIA DSX MaxLPS: NVIDIA is promoting its DSX MaxLPS technology, claiming 40% more token throughput per megawatt on Vera Rubin platforms, aiming to address energy efficiency concerns in large-scale inference clusters.
  • Speculative Decoding Adoption: Serving engines like vLLM and SGLang continue to integrate speculative decoding features, which can offer 2–5x latency improvements at the serving layer, becoming a standard expectation for efficient inference stacks.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Vera Rubin compare to Blackwell on cost?
  • QWhat drove the 36 percent Blackwell throughput gain?
  • QHow are other cloud providers matching DeepSeek?
  • QWhat are the power requirements for Jetson Thor?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.