Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-10-02
DeepSeek open-sources Ascend platform components to lower domestic AI inference costs, while CoreWeave reports higher Blackwell throughput and lower cost-per-token than traditional cloud providers. NVIDIA's Vera Rubin NVL72 posts first peer-reviewed MLPerf v6.1 results, outperforming AMD's GB300 by up to 3.7×, as software optimization alone yields 8–36% throughput gains across deployed hardware.
Inference Efficiency: MLPerf, Tokens per Dollar, Hardware — 2026-10-02
Top developments
DeepSeek Open-Sources Ascend AI Stack to Cut Domestic Inference Costs
DeepSeek announced the open-source release of a complete AI foundation component set for Huawei's Ascend platform on October 1–2, 2026, including TileLang programming language and five compute and communication libraries that map directly to NVIDIA CUDA equivalents. The dual-partner effort with Huawei targets the Ascend 950 128-GPU supernode architecture, enabling developers to reuse code across platforms and reducing migration costs for large model deployment to domestic hardware.

CoreWeave Cuts Inference Cost-per-Token Below Hyperscaler Baselines on Blackwell
CoreWeave, an AI-optimized cloud provider, delivered higher throughput and lower cost per token on NVIDIA Blackwell hardware than traditional hyperscaler offerings as of October 1, 2026. The result underscores the divergence between general-purpose cloud economics and inference-optimized infrastructure, positioning specialized providers as cost leaders in the inference-revenue-dominated AI factory model.

NVIDIA Vera Rubin NVL72 Posts 3.7× Speedup Over AMD GB300 in MLPerf v6.1
NVIDIA's Vera Rubin NVL72 achieved its first peer-reviewed MLPerf Inference v6.1 scores on September 16, 2026, beating AMD's GB300 by up to 3.7× in throughput benchmarks. The comparison marks the first official peer-reviewed performance data for Vera Rubin and reflects the expanding GPU rivalry in datacenter inference workloads.

Software Optimization Alone Yields 8–36% Throughput Gains on Deployed Blackwell
Three independent cloud providers reported 8 to 36 percent throughput improvements on already-deployed NVIDIA Blackwell hardware through software optimization alone—no new silicon required—in MLPerf v6.1 results published September 16, 2026. The finding demonstrates the ceiling for inference efficiency gains through framework and kernel tuning, with vLLM, SGLang, and TensorRT-LLM all shipping continuous batching, paged KV cache, quantization, and speculative decoding out of the box. Speculative decoding alone can deliver up to 3.6× faster generation under medium-to-low QPS workloads.
Cisco Demonstrates First Heterogeneous Multi-Vendor GPU MLPerf Submission
Cisco submitted the industry's first MLPerf Inference benchmarking results across mixed-vendor GPU pools on September 25, 2026, showing workload orchestration across heterogeneous clusters without performance penalties. The result signals enterprise interest in GPU cost optimization and vendor neutrality in inference deployments.

Local view
China: Sina and Weibo coverage emphasizes DeepSeek's strategic push to reduce "naked" (unsupported) operation of domestic AI chips on Ascend hardware. Chinese AI practitioners view the Ascend toolkit release as a critical supply-chain independence move, with CSDN and Zhihu users comparing token costs across models—DeepSeek Flash quoted at ¥1 (≈$0.14 USD) per million tokens in off-peak hours.
US: Signal65 and SiliconANGLE frame CoreWeave's cost leadership as evidence that inference economics now favor specialized cloud over hyperscalers, tied to NVIDIA's broader messaging on "AI factory economics" and inference-as-revenue. IEEE Spectrum contextualizes 2026 inference hardware as a "memory-centric" revolution splitting workloads across multiple chips.
Context & numbers
- MLPerf v6.1 participation: 30 submitters, 120 systems (record breadth) as of September 16, 2026
- Vera Rubin vs. GB300: 3.7× max throughput advantage in official benchmarks
- Software gains on Blackwell: 8–36% throughput improvement without hardware changes
- Speculative decoding speedup: Up to 3.6× faster generation under QPS-constrained inference
- DeepSeek Flash off-peak: ¥1 per million tokens (~$0.14 USD) in spare capacity
- Ascend 950 target: 128-GPU supernode architecture with TileLang/CUDA parity
On the radar
- TPUv7 Ironwood inference externalization: InferenceX preview third-party TPU benchmarks against B200/B300 with FP8 native stack (TorchTPU replacing TorchAX); watch for official Google third-party program announcement
- Agentic inference benchmarks: MLPerf Inference introduced multi-turn, closed-loop agent workload tests featuring Kimi K2.6 and Qwen 35B-A3B; token consumption 10–100× higher than single-turn for agent tasks per Futurum Research
- MLPerf Training v7 post-training round: Post-training-focused benchmarks emerging alongside inference focus; monitor MLCommons roadmap for October submission deadlines
- Cisco heterogeneous GPU pool scaling: Watch for AMD MI350X and NVIDIA Blackwell co-scheduling results in next MLPerf round (Q1 2027 expected)
Data freshness: All sources verified from October 1–2, 2026 (past 24 hours) or September 16–26, 2026. Articles older than September 25 excluded per editorial policy.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.