Efficient Training: MoE, Distillation, Compute Trends — 2026-09-10
Federal agencies issued a joint advisory identifying six AI companies involved in malicious distillation campaigns, while DeepSeek announced a significant price cut for its Flash models. Meanwhile, new research highlights that post-trained Mixture-of-Experts models can reduce expert FLOPs by over 50% via self-distillation without significant accuracy loss.
Efficient Training: MoE, Distillation, Compute Trends — 2026-09-10
Top developments
US Agencies Warn of Malicious AI Distillation Campaigns
On September 9, 2026, the NSA, CISA, and FBI issued advisory AA26-251A, explicitly naming six AI companies involved in industrial-scale extraction of capabilities from American models. The advisory details how China-based entities are using knowledge distillation to siphon intellectual property, prompting recommendations for enhanced API abuse monitoring and provenance checks. This marks a significant escalation in the security dimension of training efficiency, shifting distillation from a purely technical optimization to a compliance and security risk for model buyers.

Post-Trained MoE Can Skip Half Experts via Self-Distillation
New research titled "Post-Trained MoE Can Skip Half Experts via Self-Distillation" (ZEDA) demonstrates that post-trained Mixture-of-Experts models can eliminate over 50% of expert FLOPs with marginal accuracy loss. Tested on Qwen3-30B-A3B and GLM-4.7-Flash across 11 benchmarks, ZEDA outperformed the strongest dynamic MoE baselines by 4.0 to 6.1 points. This technique offers a concrete path to reducing inference compute costs in sparse architectures without retraining from scratch, directly impacting the efficiency of deployed MoE models.
DeepSeek Flash Models Price Cut Announced
On September 10, 2026, DeepSeek announced a full price reduction for its Flash series models, with idle-time calls dropping to 1 yuan per million tokens. This aggressive pricing strategy aims to capture market share and pressure competitors by leveraging their efficient MoE architecture to lower serving costs. The move signals a continued trend where architectural efficiency gains are passed down to API prices, intensifying the cost war among frontier model providers.
Local view
Chinese media outlets like Sohu highlighted the "price nuclear explosion" caused by DeepSeek's September 10 announcement, framing it as a direct benefit to developers and a strategic move to dominate the low-cost AI market. Meanwhile, reports from Bannedbook and others noted the US advisory naming six Chinese firms, reflecting a growing tension between US security concerns and Chinese AI industry expansion.
Context & numbers
Epoch AI data indicates that spending on training large-scale ML models is growing at a rate of 2.4x per year, with the most advanced models now costing hundreds of millions of dollars. Frontier pretraining budgets crossed the half-billion-dollar mark in 2025 and are projected to reach $1–3 billion per model by 2027. Compute (GPU-hours) remains the dominant cost line item at 65–75% of total training expenses.
On the radar
- AI Spend Trends: TechCrunch reported that AI spend per employee slumped at top firms in August 2026, raising questions about whether this is a seasonal dip or a warning sign for hyperscalers.
- Distillation Compliance: Model buyers should anticipate increased scrutiny on provenance checks and API usage patterns following the AA26-251A advisory.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.