CrewCrew
FeedSignalsMy Subscriptions
Get Started
Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-13

  1. Signals
  2. /
  3. Efficient Training: MoE, Distillation, Compute Trends

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-13

Efficient Training: MoE, Distillation, Compute Trends|September 13, 2026(1h ago)2 min read8.7AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

This week, the industry witnessed a significant pivot in training economics with the release of the Instella-MoE technical report, demonstrating high-efficiency Mixture-of-Experts training on AMD hardware. Simultaneously, US agencies issued a critical advisory on malicious AI distillation, highlighting the security risks of model extraction. Meanwhile, Chinese media reports indicate a continued downward trend in inference costs, driven by aggressive price cuts from major labs like Zhipu and DeepSeek.

Efficient Training: MoE, Distillation, Compute Trends — 2026-09-13


Top developments


Instella-MoE Demonstrates Efficient Training on AMD Instinct GPUs

The technical report for Instella-MoE-16B-A3B, published this week, details the successful pretraining of a fully open Mixture-of-Experts (MoE) language model entirely on AMD Instinct MI300X/MI325X GPUs. The model features 16 billion total parameters but only 2.8 billion active parameters per token, trained from scratch on 7.1 trillion tokens. This development is significant for compute economics as it provides a verified roadmap for achieving frontier-level MoE performance without relying exclusively on NVIDIA H100/H200 clusters, potentially diversifying the supply chain for large-scale training runs.

Instella-MoE Technical Report
Instella-MoE Technical Report


US Agencies Issue Advisory on Malicious AI Distillation

In a coordinated move, the NSA, CISA, and FBI released advisory AA26-251A last week, explicitly naming six AI companies involved in or targeted by malicious distillation campaigns. The advisory urges AI model buyers to implement stricter provenance checks and API abuse monitoring to prevent competitors from extracting proprietary capabilities through "distillation attacks." This marks a shift in how distillation is viewed: no longer just a technique for efficiency, but a vector for intellectual property theft and national security risk. For training researchers, this implies a need for robust defenses against output-based extraction methods.

US Distillation Advisory
US Distillation Advisory

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com

digitalapplied.com

six AI companies in advisory AA26-251A. See what changes for model procurement, provenance checks an

digitalapplied.com

digitalapplied.com


Inference Costs Drop as Labs Aggressively Cut Prices

Chinese tech outlet 80aj.com reported this week that the "distillation war" and efficiency gains are driving prices down further. Specifically, Zhipu’s new GLM-5.3-Flash model is priced at approximately 0.13 RMB per million tokens (approx. $0.018 USD) under limited-time discounts, undercutting DeepSeek-V4-Flash. This aggressive pricing strategy by Chinese labs is forcing global competitors to optimize their inference stacks rapidly to maintain margin viability, accelerating the adoption of quantization and sparse attention mechanisms in production environments.


Local view

Local stakeholders in China are closely monitoring the "price war" among large language models. Zhihu discussions highlight that GLM-5.3-Flash's new pricing structure makes it one of the most cost-effective options for enterprise developers, particularly for high-volume, low-latency tasks. The consensus among local analysts is that the combination of MoE architectures and distillation techniques has lowered the barrier to entry for high-quality AI services, shifting competition from raw capability to cost-per-token efficiency.


Context & numbers

  • Active Parameters: Instella-MoE uses 2.8B active parameters out of 16B total, maintaining high efficiency.
  • Training Volume: The Instella model was trained on 7.1T tokens.
  • Hardware: Instella-MoE was trained on AMD Instinct MI300X/MI325X GPUs.
  • Pricing: GLM-5.3-Flash is listed at ~0.13 RMB/M tokens during promotional periods.

On the radar

  • FP8/FP4 Adoption: The recent Megatron-Core technical report highlights Section 5 on Reduced-Precision Training in FP8/FP4 for MoE, suggesting that next-generation efficiency gains will come from lower-precision formats rather than just architectural changes.
  • Provenance Standards: Following the AA26-251A advisory, expect increased demand for tools that verify the provenance of training data and model weights to comply with new security guidelines.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Instella-MoE compare to Nvidia clusters?
  • QWhat are the six companies named in the NSA advisory?
  • QHow are labs defending against distillation attacks?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.