Efficient Training: MoE, Distillation, Compute Trends — 2026-09-27
This week's dominant story is DeepSeek's reported shift to Huawei Ascend chips for training a much larger MoE model, alongside a new DeepSeek agent-training infrastructure paper. Separately, a fresh census of lab-disclosed training costs underscores how opaque frontier compute economics remain. Compute-trend benchmarks (Epoch AI) continue to frame the escalation.
Efficient Training: MoE, Distillation, Compute Trends — 2026-09-27
Top developments
DeepSeek reportedly shifting frontier training to Huawei Ascend chips
On September 23, multiple outlets citing The Information reported that DeepSeek founder Liang Wenfeng told investors at a closed-door meeting that the company is moving large-model training to domestic chips, planning large-scale use of Huawei Ascend, with Huawei possibly beginning training-grade chip deliveries in Q4 2026. The same reports say DeepSeek is planning a model scaled to 8 trillion parameters, signaling that the MoE parameter race is entering a new tier despite hardware constraints.

DeepSeek publishes agent-training infrastructure paper
On September 24, Chinese media reported DeepSeek disclosed a new paper on agent (RL) training infrastructure, with 130+ authors and Liang Wenfeng as the final author, describing a single cluster running 3 million sandbox environments per day and anti-"reward-hacking" measures against AI cheating. It matters for compute economics because RL/agent training, not just pretraining, is becoming a major compute sink.
Training-cost disclosure census: only 6 of 21 self-reported figures include dollar amounts
A September 22 analysis of 21 training-cost figures labs published themselves between 2022 and 2026 found only six state a dollar amount, none is independently audited, and each omits key cost components. For training-cost research, this reinforces reliance on third-party estimates and amortized cluster-cost methodologies.
Distillation re-enters the geopolitical spotlight
On September 25, the New York Times reported that as Xi Jinping visits Washington, leaders were expected to discuss claims that China is copying American AI technologies via distillation. The distillation debate directly affects how frontier labs guard API outputs, a key input for efficient smaller-model training.

Chinese coverage frames the MoE "parameter race"
HuSHARE-linked coverage (Huxiu, Sept 22) argued the industry's large-parameter race is now migrating to MoE architectures — the architectural pattern DeepSeek's reported 8T-parameter plan would follow.
Local view
Chinese-language analysis (Zhihu, Sept 23) is focused on post-"Coding Plan" model price comparisons, noting Zhipu's GLM-5.3-Flash at ~0.13 RMB/M tokens (50% off) undercutting deepseek-v4-flash — evidence that Chinese efficient-training economics are flowing straight into aggressive inference pricing. Sina's ML digest (Sept 23) highlighted McKinsey warnings that multi-step AI agents can vary up to 30x in cost for the same task, pushing enterprise spend scrutiny.

Context & numbers
- Epoch AI: frontier training compute has grown ~0.7 orders of magnitude per year since 2020; large-scale training spending is growing ~2.4x per year, with training costs doubling roughly every eight months for the largest models.
- Third-party estimates put 2026 frontier runs at 1e26–1e27 FLOPs with named cost ranges of $200–500M; for a ~$2B run, GPU-hours represent 65–75% of spend, with failed experiments adding 20–30% on top.
- Anchor model disclosures for comparison: DeepSeek-V3 (671B total / 37B active MoE, MLA, FP8, 2.788M H800 GPU-hours, $5.576M).
- Largest known AI data center: ~1.1 million (GPU-equivalent) computing capacity; gigawatt-scale data centers take about 2 years to build.
On the radar
- Watch for Huawei Ascend training-chip deliveries to DeepSeek, reportedly possible as early as Q4 2026.
- DeepSeek's next flagship technical report — whether it confirms an 8T-parameter MoE run and how much is trained on Ascend rather than NVIDIA.
- Distillation topic expected at the highest diplomatic level following the Xi–Trump meeting; policy responses to "adversarial distillation" remain under discussion.
- Rumor flag: third-party frontier-cost ranges ($200M–$2B+) are estimates, not disclosed figures — treat with caution pending lab disclosures.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.