CrewCrew
FeedSignalsMy Subscriptions
Get Started
Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-11

  1. Signals
  2. /
  3. Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-11

Efficient Training: MoE, Distillation, Compute Trends|September 11, 2026(5h ago)3 min read9.1AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

This week, the focus on AI training efficiency shifted from pure technical optimization to geopolitical and economic security, with US agencies issuing a major advisory on "malicious distillation" by Chinese firms. Simultaneously, DeepSeek announced significant price cuts for its Flash models, signaling a new phase in compute economics where inference costs are being aggressively reduced to drive adoption.

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-11


Top developments


US Agencies Issue Advisory on "Malicious Distillation"

On September 9, 2026, the NSA, CISA, and FBI jointly released Advisory AA26-251A, explicitly naming six Chinese AI companies and warning that they are engaging in "industrial-scale extraction" of capabilities from leading American models. The advisory details how these firms use large-scale knowledge distillation—querying proprietary APIs with millions of prompts to generate training data—to replicate advanced reasoning capabilities without paying for the original compute or data. This move marks a significant escalation in the "AI wars," shifting the definition of IP theft from code copying to behavioral replication via API abuse.

US CISA logo representing the recent advisory on AI distillation
US CISA logo representing the recent advisory on AI distillation

helpnetsecurity.com

helpnetsecurity.com


DeepSeek Announces Aggressive Price Cuts for Flash Models

On September 10, 2026, DeepSeek officially launched a massive price reduction for its Flash series models, cutting costs by up to 60% in certain tiers. The new off-peak pricing is set at 1 RMB per million tokens for input (cache miss) and 4 RMB per million tokens for output, with cached inputs dropping to 0.02 RMB. This aggressive pricing strategy leverages their efficient Mixture-of-Experts (MoE) architecture to undercut competitors, aiming to capture market share in the inference layer despite rising upstream compute costs.

DeepSeek logo illustrating the recent pricing updates
DeepSeek logo illustrating the recent pricing updates


TechCrunch Reports Slump in Per-Employee AI Spend

Data released on September 9, 2026, indicates that AI spending per employee at top tech firms slumped in August. The report attributes this decline to falling token costs and the availability of cheaper, more efficient models. While some view this as a seasonal dip, analysts warn it could signal a "warning sign" for hyperscalers who have bet on ever-increasing per-user consumption, suggesting that efficiency gains are finally translating into lower enterprise bills rather than just higher margins.


Local view

In Chinese-language media, the narrative surrounding the US advisory is one of defensive posturing against "technological containment." Outlets like Bannedbook (citing VOA) report that the US government is targeting companies like DeepSeek and Moonshot AI specifically because their models have achieved near-parity with US frontier models at a fraction of the cost. Meanwhile, financial media like Investment Circle (Touzi Jie) frame DeepSeek’s price cuts not just as a commercial tactic, but as a strategic move to solidify the domestic ecosystem's independence from expensive foreign APIs, highlighting the "price explosion" as a win for local developers.


Context & numbers

The backdrop to these events is a rapidly changing compute landscape. Frontier training costs are projected to hit $1–3 billion per model by 2027, with power availability becoming a tighter constraint than GPU supply. However, the cost per unit of intelligence is dropping. Epoch AI data trends show that while total spend grows, the efficiency of MoE architectures allows for significant performance gains without proportional increases in active parameter counts. For instance, DeepSeek’s earlier reports indicated training efficiency metrics (MFU) that rivaled much larger, more expensive US models, proving that architectural innovation can offset raw compute disadvantages.


On the radar

  • DeepSeek IPO Preparation: Reports confirm DeepSeek has hired CITIC Securities to prepare for a Science and Technology Innovation Board (STAR Market) IPO, potentially launching this year. This could provide massive capital for future compute clusters.
  • Legal Frameworks for Distillation: Following the US advisory, legal experts are debating how to define "illegal" distillation versus standard model training, which may lead to new API terms of service enforcement or regulatory interventions in Q4 2026.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow will US agencies stop malicious distillation?
  • QHow do competitors respond to DeepSeek's price cuts?
  • QAre falling enterprise AI bills a trend for good?
  • QWhat Chinese companies are named in the advisory?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.