Efficient Training: MoE, Distillation, Compute Trends — 2026-09-12
This week, the US government formally named six AI companies in an advisory regarding "industrialized extraction" of model capabilities, intensifying the legal and security debate around model distillation. Meanwhile, DeepSeek executed a significant API price reduction, signaling a shift in compute economics, while new technical reports from AMD-backed research highlighted efficient MoE training on non-NVIDIA hardware.
Efficient Training: MoE, Distillation, Compute Trends — 2026-09-12
Top developments
US Advisory Targets "Industrialized" Distillation by Six AI Firms
On September 5, 2026, the NSA, CISA, and FBI issued advisory AA26-251A, explicitly naming six AI companies involved in the "industrialized extraction" of capabilities from US models. This advisory marks a significant escalation from general warnings to specific attribution, urging model buyers to implement stricter provenance checks and API abuse monitoring to detect malicious distillation attempts.
Anthropic Alleges 185M+ Responses Extracted in Distillation Attacks
Anthropic released a threat intelligence report on September 11, 2026, alleging that seven China-based labs conducted large-scale distillation attacks against Claude, extracting over 185 million responses. The report details how these actors used automated pipelines to query the model and train competing systems, raising urgent questions about the legal recourse available to frontier labs and the security of API-based training data.

DeepSeek Cuts Prices by Up to 60% to Drive Adoption
On September 10, 2026, DeepSeek announced a major price reduction for its flash series models, with idle-time prices for cache hits dropping to ¥0.02 per million tokens (approx. $0.003), representing a cut of up to 60% from previous rates. This aggressive pricing strategy aims to recapture market share lost during its earlier peak/off-peak pricing introduction, forcing competitors to reconsider their own cost structures and efficiency gains.
Instella-MoE Demonstrates Efficient Training on AMD Hardware
A new technical report for Instella-MoE-16B-A3B was published around September 8, 2026, detailing the training of a 16B parameter Mixture-of-Experts model with only 2.8B active parameters per token. Trained from scratch on 7.1T tokens using AMD Instinct MI300X/MI325X GPUs, the project provides a complete open-source pipeline, challenging NVIDIA's dominance in efficient large-scale training infrastructure.

Local view
Chinese tech media reported extensively on DeepSeek's price cuts, with Pedaily (Investment Community) highlighting the "return of the King Liang" meme among developers as prices dropped to historic lows. Meanwhile, 80aj noted in its weekly AI digest that while domestic models are engaging in price wars, the international focus has shifted sharply toward security advisories regarding distillation, creating a divergent narrative between Chinese cost-efficiency achievements and US security concerns.
Context & numbers
- Distillation Volume: Anthropic reports over 185 million responses were extracted via automated queries in recent distillation campaigns.
- DeepSeek Pricing: Idle-time input prices for DeepSeek Flash dropped to ¥0.02/million tokens (cache hit) and ¥4/million tokens (output) effective September 10, 2026.
- MoE Efficiency: The Instella-MoE model activates only 2.8B parameters out of 16B total, demonstrating high sparsity efficiency on AMD hardware.
On the radar
- Legal Fallout: Watch for initial legal filings or sanctions discussions following the US advisory AA26-251A naming six specific companies.
- Competitor Price Responses: Major US and Chinese labs may announce counter-pricing or efficiency claims in response to DeepSeek's latest cuts.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.