CrewCrew
FeedSignalsMy Subscriptions
Get Started
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-28

  1. Signals
  2. /
  3. Inference Efficiency: MLPerf, Tokens per Dollar, Hardware

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-28

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware|September 28, 2026(2h ago)3 min read8.3AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

This week's news is dominated by pricing rather than benchmarks: Anthropic and OpenAI cut frontier inference prices within hours of each other, and an Epoch report argues fixed-capability inference cost falls roughly 13x per year. Meanwhile, Chinese providers continue aggressive per-million-token discounting, and DeepSeek's move toward "ultra-large" parameters signals rising serving costs despite falling unit prices.

Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-09-28


Top developments

Source image
Source image

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

mlcommons.org

MLCommons Sets Participation Record with New MLPerf Inference v6.1 Benchmark Results - MLCommons

mlcommons.org

Where the Industry Is Investing: A Look at MLPerf Inference v6.1 - MLCommons


Anthropic and OpenAI kick off the inference price war

Within hours of each other this week, Anthropic and OpenAI cut inference prices, with Value Add Pulse tallying the current per-token landscape: Opus 5.5 at $4/$20 per 1M tokens versus GPT-6 Sol at $2/$10 per 1M. The back-to-back cuts show frontier labs competing directly on price rather than deferring to capability differentiation. This matters for tokens-per-dollar math: headline benchmark throughput matters less than what providers actually charge per delivered token.

Source image
Source image

lambda.ai

lambda.ai


Epoch: fixed-capability inference cost falls ~13x per year

A report covered by Huxiu argues that AI inference cost at fixed capability declines roughly 13x annually, framing "intelligence" as becoming a priced, metered infrastructure commodity. This rate of decline is faster than most enterprise budgeting assumptions and helps explain how vendors can repeatedly cut prices while still running newer, larger models on the same hardware class.


DeepSeek joins the parameter race — and that raises serving bills

21jingji reports DeepSeek is entering the ultra-large parameter race, with the headline citing an "8 trillion" parameter figure as the company, led by Liang Wenfeng, pushes scale. Larger models mean higher serving costs per token unless aggressively optimized, which contextualizes DeepSeek's recent flash-tier discounting: cheap off-peak pricing on small models subsidizes headline-grabbing large-model development.


Agentic workloads may consume 10–100x more tokens per task

Futurum Research, cited in jCodeMunch's September 26 token-cost radar, estimates agentic AI consumes 10 to 100 times more tokens per task than single-pass inference. The piece argues the real story this week is "challenging the meter itself" — per-token pricing may not survive agentic workflows, pushing the industry toward per-outcome or per-task billing.


Local view

Chinese developer communities are intensely focused on per-million-token price floors. A Zhihu comparison notes Zhipu's GLM-5.3-Flash, at a limited-time 50% discount, costs roughly ¥0.13 per 1M tokens pay-as-you-go — slightly below DeepSeek-v4-flash — with lite idle pricing around ¥0.09 and peak at ¥0.19, or about one-third of GLM-5.3 plan pricing. Also notable for Chinese readers: a V2EX tracker shows DeepSeek-v4.1-flash deployment and pricing across domestic platforms, with Alibaba Qianwen live but priced roughly 50% above official rates, and Baidu Qianfan and Kuaishou Wanqing not yet online.


Context & numbers

  • Opus 5.5: $4 input / $20 output per 1M tokens; GPT-6 Sol: $2/$10 per 1M
  • Epoch estimate: fixed-capability inference cost declining ~13x per year
  • Agentic tasks: 10–100x more tokens per task than standard inference, per Futurum Research
  • GLM-5.3-Flash: ~¥0.13/1M token (discounted), lite idle ~¥0.09, peak ~¥0.19
  • CloudZero finds the same open model can cost 9x more on one inference provider than another, ranking 16+ providers on cost per outcome rather than per token

On the radar

  • Gemini 3.8 Flash prices are slated to double on January 1, 2027, per the beri.net cost breakdown — worth watching for pre-deadline buying behavior (flagged as reported pricing, not confirmed policy).
  • CloudZero's provider ranking on "cost per outcome" may pressure vendors to publish task-level, not just per-token, pricing.
  • Watch whether the next MLPerf Inference round adopts agent-heavy workloads at scale, given that agentic token consumption of 10–100x is becoming the dominant cost driver (rumor/speculation based on current benchmark direction).

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow will agentic workflows change AI billing models?
  • QWhat hardware drives the 13x annual cost decline?
  • QHow are enterprises adjusting to lower token costs?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.