Efficient Training: MoE, Distillation, Compute Trends — 2026-09-23
The big news this week is DeepSeek's reported plan to scale to a 2 trillion-parameter model next and 8 trillion after, cementing MoE as the arena for the new parameter race even as rivals cut API prices. Frontier launches from OpenAI and Anthropic turned the model race into a price war, while Stanford's AI Index findings recirculated: training compute doubles roughly every five months. A new practitioner guide argues task-specific distillation can deliver production models at ~1% of frontier cost.
Efficient Training: MoE, Distillation, Compute Trends — 2026-09-23
Top developments
DeepSeek reportedly training 2T-parameter model, 8T planned
According to The Information, DeepSeek CEO Liang Wenfeng told investors the company is training a 2-trillion-parameter model and plans an 8-trillion-parameter one after; the current flagship V4 Pro has 1.6T total parameters. For comparison, the largest publicly released supermodel, Moonshot AI's Kimi K3, reaches 2.8T parameters. The move confirms that MoE sparsity — not dense scaling — is the route to very large parameter counts, with direct implications for compute economics per activated token.

Frontier launches turn the model race into a price war
On Tuesday, September 22, OpenAI launched GPT-6 Sol and Luna while Anthropic shipped Claude Opus 5.5, with pricing pressure emerging as the competitive axis; Meta Muse and Alibaba's chip efforts also featured in the day's news. For training-efficiency watchers, simultaneous flagship launches signal that headline capability is no longer differentiating — cost-per-token is, which rewards labs with cheaper training pipelines (MoE, low-precision, distillation).
Stanford AI Index: training compute doubling every 5 months
Coverage recirculating this week highlights Stanford's AI Index findings that training compute for notable AI models doubles roughly every five months, LLM dataset sizes every eight months, and the electrical power needed for training roughly every year — a pace about four times Moore's Law. This is the baseline arithmetic behind the push into FP8 training, MoE sparsity, and synthetic data: efficiency gains must outrun a brutal compute curve.
Distillation as a production strategy: 1% of frontier cost
A guide published this week frames task-specific student models distilled from frontier teachers as costing roughly 1% of the frontier model when deployed in production, and lays out when distillation beats quantization, RAG, or paying the API bill. It matters because distillation is becoming a standard compute-cost hedge for enterprises as frontier API pricing shifts.

Local view
Chinese-language coverage is focused on the parameter race and its architecture. Huxiu framed the reported 8T-parameter plan as evidence that the large-model competition has pivoted to MoE architectures ("大模型参数竞赛转向MoE架构"). A research note from Orient Securities (东方证券), relayed via Ifeng, argued DeepSeek's V4.1 Flash at only 552B parameters achieves strong benchmark results with less compute — cutting HBM requirements to 1/4 and SSD needs to 1/8 versus the prior generation — and expects this to accelerate domestic compute demand. Zhihu posters are tracking a post-"Coding Plan" era of model price comparisons, with Zhipu's GLM-5.3-Flash at a limited-time rate of about ¥0.13 per million tokens, slightly below DeepSeek's v4-flash.
Context & numbers
- Frontier training runs in 2026 sit between 1e26 and 1e27 FLOP, with named cost ranges of $200–500M for the GPT-5/Gemini Ultra class and projections of $1–3B for the late-2027 frontier; compute (GPU-hours) is 65–75% of a ~$2B run's budget, with failed experiments adding 20–30%.
- Epoch AI: frontier training compute has grown ~0.7 orders of magnitude per year since 2020; the largest AI data center has capacity equivalent to 1.1 million (H100-class) chips, and gigawatt-scale facilities take ~2 years to build.
- The AI training chip market is projected to grow from $8.44B in 2025 to $29.4B by 2033 (16.9% CAGR).
- Reference architecture numbers: DeepSeek-V3 uses 671B total parameters with 37B activated per token; FP8 storage of activations saves ~16 GB (~12% of its 131 GB activation budget) per the Megatron Core MoE technical report.
On the radar
- Watch for DeepSeek V4.1 Pro's official technical report — reported 2T-parameter scale and claimed domestic-chip training will be scrutinized for FP8 usage and cost disclosures (rumor stage until confirmed).
- DeepSeek suffered its 18th large-scale outage of the year on September 20 (web, app, and API down simultaneously) — worth monitoring for infrastructure reliability as training/inference loads scale.
- The Instella-MoE-16B-A3B fully open MoE model (2.8B active, trained on AMD MI300X/MI325X) continues to circulate as a reproducibility reference for non-NVIDIA MoE training.
- Enterprise AI-compute cost coverage is climbing (e.g., NetNewsLedger on how per-million-token pricing distorts real spend) — expect more scrutiny of how training efficiency translates to list pricing.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.