CrewCrew
FeedSignalsMy Subscriptions
Get Started
Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-12

  1. Signals
  2. /
  3. Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-12

Efficient Training: MoE, Distillation, Compute Trends|September 12, 2026(2h ago)2 min read8.7AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

This week, the US government formally named six AI companies in an advisory regarding "industrialized extraction" of model capabilities, intensifying the legal and security debate around model distillation. Meanwhile, DeepSeek executed a significant API price reduction, signaling a shift in compute economics, while new technical reports from AMD-backed research highlighted efficient MoE training on non-NVIDIA hardware.

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-12


Top developments


US Advisory Targets "Industrialized" Distillation by Six AI Firms

On September 5, 2026, the NSA, CISA, and FBI issued advisory AA26-251A, explicitly naming six AI companies involved in the "industrialized extraction" of capabilities from US models. This advisory marks a significant escalation from general warnings to specific attribution, urging model buyers to implement stricter provenance checks and API abuse monitoring to detect malicious distillation attempts.

US Distillation Advisory graphic
US Distillation Advisory graphic

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com

six AI companies in advisory AA26-251A. See what changes for model procurement, provenance checks an

digitalapplied.com

digitalapplied.com


Anthropic Alleges 185M+ Responses Extracted in Distillation Attacks

Anthropic released a threat intelligence report on September 11, 2026, alleging that seven China-based labs conducted large-scale distillation attacks against Claude, extracting over 185 million responses. The report details how these actors used automated pipelines to query the model and train competing systems, raising urgent questions about the legal recourse available to frontier labs and the security of API-based training data.

AI Distillation Attacks 2026 graphic
AI Distillation Attacks 2026 graphic

ailearningguides.com

ailearningguides.com


DeepSeek Cuts Prices by Up to 60% to Drive Adoption

On September 10, 2026, DeepSeek announced a major price reduction for its flash series models, with idle-time prices for cache hits dropping to ¥0.02 per million tokens (approx. $0.003), representing a cut of up to 60% from previous rates. This aggressive pricing strategy aims to recapture market share lost during its earlier peak/off-peak pricing introduction, forcing competitors to reconsider their own cost structures and efficiency gains.


Instella-MoE Demonstrates Efficient Training on AMD Hardware

A new technical report for Instella-MoE-16B-A3B was published around September 8, 2026, detailing the training of a 16B parameter Mixture-of-Experts model with only 2.8B active parameters per token. Trained from scratch on 7.1T tokens using AMD Instinct MI300X/MI325X GPUs, the project provides a complete open-source pipeline, challenging NVIDIA's dominance in efficient large-scale training infrastructure.

Instella-MoE paper thumbnail
Instella-MoE paper thumbnail


Local view

Chinese tech media reported extensively on DeepSeek's price cuts, with Pedaily (Investment Community) highlighting the "return of the King Liang" meme among developers as prices dropped to historic lows. Meanwhile, 80aj noted in its weekly AI digest that while domestic models are engaging in price wars, the international focus has shifted sharply toward security advisories regarding distillation, creating a divergent narrative between Chinese cost-efficiency achievements and US security concerns.


Context & numbers

  • Distillation Volume: Anthropic reports over 185 million responses were extracted via automated queries in recent distillation campaigns.
  • DeepSeek Pricing: Idle-time input prices for DeepSeek Flash dropped to ¥0.02/million tokens (cache hit) and ¥4/million tokens (output) effective September 10, 2026.
  • MoE Efficiency: The Instella-MoE model activates only 2.8B parameters out of 16B total, demonstrating high sparsity efficiency on AMD hardware.

On the radar

  • Legal Fallout: Watch for initial legal filings or sanctions discussions following the US advisory AA26-251A naming six specific companies.
  • Competitor Price Responses: Major US and Chinese labs may announce counter-pricing or efficiency claims in response to DeepSeek's latest cuts.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QWhich six AI firms were named in the US advisory?
  • QHow will Anthropic legally respond to the attacks?
  • QHow are competitors reacting to DeepSeek's price cuts?
  • QWhat is the performance of AMD hardware vs NVIDIA?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.