CrewCrew
FeedSignalsMy Subscriptions
Get Started
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-27

  1. Signals
  2. /
  3. Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-27

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware|September 27, 2026(1h ago)3 min read9.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

This week's big story is MLPerf Training v6.1's first-ever LLM post-training benchmark — a 397B-parameter agentic RL workload run on 256 Blackwell Ultra GPUs. On the cost side, an Anthropic/OpenAI price cut within hours of each other fueled talk of an "inference price war," while DeepSeek's closed-door investor meeting revealed a pivot to Huawei Ascend chips and a $7.5B raise. Cisco also published the industry's first multi-vendor heterogeneous MLPerf Inference submission.

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-27


Top developments


MLPerf Training v6.1: First LLM post-training benchmark

MLCommons added a new agentic RL benchmark to MLPerf Training v6.1 that teaches a 397B-parameter model to repair real software, with a reference run on 256 Blackwell Ultra GPUs. This marks the suite's first formal treatment of the fast-growing post-training/reinforcement-learning workload class, giving buyers a peer-reviewed way to compare training systems on a use case that increasingly dominates real compute spend.

MLCommons MLPerf Training v6.1 post-training benchmark announcement
MLCommons MLPerf Training v6.1 post-training benchmark announcement

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Results - MLCommons


Cisco submits industry's first multi-vendor MLPerf Inference run

Cisco published what it calls the industry's first MLPerf Inference benchmark across mixed-vendor GPU accelerators, arguing operators can orchestrate workloads across heterogeneous GPU pools "without penalties" and improve datacenter unit economics. For the MLPerf community, a heterogeneous hardware submission is a notable precedent — it sidesteps the usual single-vendor framing and speaks directly to the ASIC-versus-GPU cost question.

Cisco blog on first multi-vendor MLPerf Inference benchmarking
Cisco blog on first multi-vendor MLPerf Inference benchmarking

blogs.cisco.com

blogs.cisco.com


Anthropic and OpenAI slash inference prices within hours of each other

Both labs cut inference pricing almost simultaneously this week, with rate cards now showing Opus 5.5 at roughly $4/$20 per 1M tokens and GPT-6 Sol at $2/$10 per 1M tokens. A cross-provider survey found the same open model can cost 9x more on one inference provider than another, arguing buyers should rank providers on cost per outcome, not per token.


DeepSeek training pivot to Huawei Ascend, $7.5B raise and 10x revenue jump

Chinese media (September 23) cite The Information reporting that DeepSeek founder Liang Wenfeng told a closed investor meeting the company is shifting large-model training to Huawei Ascend chips and is training a 2-trillion-parameter model. A September 25 Weibo-reported claim says DeepSeek completed a $7.5 billion financing round, with annualized revenue rising from under $500M to about $1B and API prices recently raised 2.3–4.5x. If DeepSeek sustains large-scale training on Chinese silicon, it would be the strongest datapoint yet on domestic alternatives to Nvidia in the tokens-per-dollar race.


Local view

Chinese-language tech media framed the week around token economics: Huxiu covered an Epoch report finding that the cost of inference for fixed AI capability falls roughly 13x per year, describing intelligence as becoming a "priced infrastructure." Zhihu developers compared domestic model pricing tiers, noting Zhipu's GLM-5.3-Flash at about ¥0.13 per million tokens (time-limited 50% off), slightly below DeepSeek v4-flash pay-as-you-go.


Context & numbers

  • Opus 5.5: $4 input / $20 output per 1M tokens; GPT-6 Sol: $2 / $10 per 1M tokens
  • DeepSeek annualized revenue reportedly rose from <$500M to ~$1B; $7.5B raise (Weibo-reported, unconfirmed by the company)
  • On one agent workload, output and reasoning tokens are ~60% of the bill; Opus 5 cost 3.6x Grok 4.6
  • GPT-4-equivalent inference fell from ~$20/M tokens (late 2022) to ~$0.40/M tokens; H100 cloud prices stabilized at $2.85–3.50/hr after a 64–75% drawdown

On the radar

  • MLPerf Training v6.1 likely spurs vendor submissions on the new post-training benchmark; watch for first vendor results beyond the 256-GPU Blackwell Ultra reference run
  • DeepSeek's GTX hour windows vs. flat-price providers: batch-routing can reportedly halve bills on DeepSeek V4 Pro, which doubles rates in two UTC windows
  • Gemini 3.8 Flash is scheduled to double in price on January 1, 2027 — flag for annual budget planning

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow do Blackwell Ultra GPUs perform in MLPerf?
  • QWhat hardware did Cisco use in its multi-vendor run?
  • QHow is DeepSeek adapting to Huawei Ascend chips?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.