CrewCrew
FeedSignalsMy Subscriptions
Get Started
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-04

  1. Signals
  2. /
  3. Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-04

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware|September 4, 2026(1h ago)3 min read8.7AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

MLCommons released MLPerf Storage v3.0 results this week, highlighting Everpure’s dominance in checkpointing and KV cache performance with speeds reaching 877 GiB/s. Meanwhile, global inference costs continued their steep decline, dropping 43% over ten weeks to an average of $1.16 per million tokens, driven by aggressive price cuts from OpenAI and Chinese providers like DeepSeek.

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-04


Top developments


MLPerf Storage v3.0 Benchmark Results Released

On September 1, 2026, MLCommons officially released the results for MLPerf Storage v3.0, a significant update to the benchmark suite that now includes tests for Key-Value (KV) cache performance and vector database workloads alongside traditional checkpointing. The new version introduces an S3 data layer option in addition to POSIX, reflecting the shift toward cloud-native AI infrastructure. This update is critical for MLPerf Inference as it provides a more holistic view of how storage bottlenecks impact large model serving and training efficiency.

MLPerf Storage v3.0 Checkpoint Scaling Chart
MLPerf Storage v3.0 Checkpoint Scaling Chart

storagereview.com

storagereview.com


Everpure Leads in Large-Model Checkpointing and KV Cache

In the newly released v3.0 results, Everpure’s FlashBlade//EXA system secured the number one ranking across multiple categories for large-model checkpointing and KV cache performance, specifically for models with 405B and 1.25T parameters. The system achieved checkpointing speeds of up to 877 GiB/s. These metrics are vital for inference efficiency because faster KV cache retrieval directly reduces latency in long-context LLM serving, a key metric in modern token-per-second benchmarks.

Everpure FlashBlade//EXA System
Everpure FlashBlade//EXA System

hpcwire.com

hpcwire.com


AI Inference Costs Drop 43% in Ten Weeks

Data from early September 2026 indicates that average AI inference costs have plummeted by 43% over the last ten weeks, reaching a record low of $1.16 per million tokens. This rapid deflation is largely attributed to OpenAI slashing prices on GPT-5.6 models by 80% and intense competition from Chinese AI labs. The drop in cost-per-token forces hardware vendors to focus more heavily on efficiency metrics like tokens per dollar rather than just raw throughput, reshaping the competitive landscape for ASICs versus GPUs.


NVIDIA and Cerebras Criticized for "Batch 1" Marketing

The Register published an analysis on August 27, 2026, criticizing NVIDIA and Cerebras for marketing peak inference speeds measured at batch size 1, which they argue does not reflect real-world customer usage patterns. The article highlights that while Cerebras claims 4,400 tokens/s and NVIDIA claims 3,400 tokens/s in single-stream scenarios, these numbers degrade significantly at higher context lengths and batch sizes required for enterprise workloads. This debate underscores the growing disconnect between vendor-reported "tokens per second" records and independent MLPerf-style server throughput benchmarks.


Local view

No recent local-language media coverage specific to this signal's niche was identified in the provided research results for the past 7 days.


Context & numbers

  • Average Cost: $1.16 per million tokens (as of early Sept 2026)
  • Top Storage Speed: 877 GiB/s checkpointing speed (Everpure FlashBlade//EXA)
  • Benchmark Submissions: 143 total submissions in MLPerf Storage v3.0

On the radar

  • DeepSeek Pricing Adjustments: Following the August 17 implementation of peak/off-peak pricing, developers are reporting significant cost increases for high-volume users, with some seeing daily bills jump from ¥1.8 to ¥9.7 for similar usage patterns due to peak-hour surcharges.
  • Speculative Decoding Adoption: Recent comparisons between vLLM, SGLang, and TensorRT-LLM continue to highlight speculative decoding as a key lever for achieving 2-5x latency improvements, with SGLang showing 15-30% higher throughput than vLLM on H100s in certain configurations.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow do GPUs and ASICs compare on tokens per dollar?
  • QWhat drove the 43% drop in AI inference costs?
  • QHow does batch size impact real-world inference speed?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.