CrewCrew
FeedSignalsMy Subscriptions
Get Started
Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-23

  1. Signals
  2. /
  3. Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-23

Efficient Training: MoE, Distillation, Compute Trends|September 23, 2026(3h ago)4 min read7.8AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

The big news this week is DeepSeek's reported plan to scale to a 2 trillion-parameter model next and 8 trillion after, cementing MoE as the arena for the new parameter race even as rivals cut API prices. Frontier launches from OpenAI and Anthropic turned the model race into a price war, while Stanford's AI Index findings recirculated: training compute doubles roughly every five months. A new practitioner guide argues task-specific distillation can deliver production models at ~1% of frontier cost.

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-23


Top developments


DeepSeek reportedly training 2T-parameter model, 8T planned

According to The Information, DeepSeek CEO Liang Wenfeng told investors the company is training a 2-trillion-parameter model and plans an 8-trillion-parameter one after; the current flagship V4 Pro has 1.6T total parameters. For comparison, the largest publicly released supermodel, Moonshot AI's Kimi K3, reaches 2.8T parameters. The move confirms that MoE sparsity — not dense scaling — is the route to very large parameter counts, with direct implications for compute economics per activated token.

Cover image of the Huxiu article on DeepSeek's plan to train an 8-trillion-parameter MoE model
Cover image of the Huxiu article on DeepSeek's plan to train an 8-trillion-parameter MoE model


Frontier launches turn the model race into a price war

On Tuesday, September 22, OpenAI launched GPT-6 Sol and Luna while Anthropic shipped Claude Opus 5.5, with pricing pressure emerging as the competitive axis; Meta Muse and Alibaba's chip efforts also featured in the day's news. For training-efficiency watchers, simultaneous flagship launches signal that headline capability is no longer differentiating — cost-per-token is, which rewards labs with cheaper training pipelines (MoE, low-precision, distillation).


Stanford AI Index: training compute doubling every 5 months

Coverage recirculating this week highlights Stanford's AI Index findings that training compute for notable AI models doubles roughly every five months, LLM dataset sizes every eight months, and the electrical power needed for training roughly every year — a pace about four times Moore's Law. This is the baseline arithmetic behind the push into FP8 training, MoE sparsity, and synthetic data: efficiency gains must outrun a brutal compute curve.


Distillation as a production strategy: 1% of frontier cost

A guide published this week frames task-specific student models distilled from frontier teachers as costing roughly 1% of the frontier model when deployed in production, and lays out when distillation beats quantization, RAG, or paying the API bill. It matters because distillation is becoming a standard compute-cost hedge for enterprises as frontier API pricing shifts.

Hero image from the LLM distillation production guide on teacher-student model pipelines
Hero image from the LLM distillation production guide on teacher-student model pipelines

cloudrps.com

cloudrps.com


Local view

Chinese-language coverage is focused on the parameter race and its architecture. Huxiu framed the reported 8T-parameter plan as evidence that the large-model competition has pivoted to MoE architectures ("大模型参数竞赛转向MoE架构"). A research note from Orient Securities (东方证券), relayed via Ifeng, argued DeepSeek's V4.1 Flash at only 552B parameters achieves strong benchmark results with less compute — cutting HBM requirements to 1/4 and SSD needs to 1/8 versus the prior generation — and expects this to accelerate domestic compute demand. Zhihu posters are tracking a post-"Coding Plan" era of model price comparisons, with Zhipu's GLM-5.3-Flash at a limited-time rate of about ¥0.13 per million tokens, slightly below DeepSeek's v4-flash.


Context & numbers

  • Frontier training runs in 2026 sit between 1e26 and 1e27 FLOP, with named cost ranges of $200–500M for the GPT-5/Gemini Ultra class and projections of $1–3B for the late-2027 frontier; compute (GPU-hours) is 65–75% of a ~$2B run's budget, with failed experiments adding 20–30%.
  • Epoch AI: frontier training compute has grown ~0.7 orders of magnitude per year since 2020; the largest AI data center has capacity equivalent to 1.1 million (H100-class) chips, and gigawatt-scale facilities take ~2 years to build.
  • The AI training chip market is projected to grow from $8.44B in 2025 to $29.4B by 2033 (16.9% CAGR).
  • Reference architecture numbers: DeepSeek-V3 uses 671B total parameters with 37B activated per token; FP8 storage of activations saves ~16 GB (~12% of its 131 GB activation budget) per the Megatron Core MoE technical report.

On the radar

  • Watch for DeepSeek V4.1 Pro's official technical report — reported 2T-parameter scale and claimed domestic-chip training will be scrutinized for FP8 usage and cost disclosures (rumor stage until confirmed).
  • DeepSeek suffered its 18th large-scale outage of the year on September 20 (web, app, and API down simultaneously) — worth monitoring for infrastructure reliability as training/inference loads scale.
  • The Instella-MoE-16B-A3B fully open MoE model (2.8B active, trained on AMD MI300X/MI325X) continues to circulate as a reproducibility reference for non-NVIDIA MoE training.
  • Enterprise AI-compute cost coverage is climbing (e.g., NetNewsLedger on how per-million-token pricing distorts real spend) — expect more scrutiny of how training efficiency translates to list pricing.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does DeepSeek fund its 8T-parameter training?
  • QWhat hardware powers GPT-6 Sol and Claude Opus 5.5?
  • QHow fast will the Stanford AI Index compute curve go?
  • QWhen should enterprises choose distillation over RAG?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.