Efficient Training: MoE, Distillation, Compute Trends — 2026-09-13
This week, the industry witnessed a significant pivot in training economics with the release of the Instella-MoE technical report, demonstrating high-efficiency Mixture-of-Experts training on AMD hardware. Simultaneously, US agencies issued a critical advisory on malicious AI distillation, highlighting the security risks of model extraction. Meanwhile, Chinese media reports indicate a continued downward trend in inference costs, driven by aggressive price cuts from major labs like Zhipu and DeepSeek.
Efficient Training: MoE, Distillation, Compute Trends — 2026-09-13
Top developments
Instella-MoE Demonstrates Efficient Training on AMD Instinct GPUs
The technical report for Instella-MoE-16B-A3B, published this week, details the successful pretraining of a fully open Mixture-of-Experts (MoE) language model entirely on AMD Instinct MI300X/MI325X GPUs. The model features 16 billion total parameters but only 2.8 billion active parameters per token, trained from scratch on 7.1 trillion tokens. This development is significant for compute economics as it provides a verified roadmap for achieving frontier-level MoE performance without relying exclusively on NVIDIA H100/H200 clusters, potentially diversifying the supply chain for large-scale training runs.

US Agencies Issue Advisory on Malicious AI Distillation
In a coordinated move, the NSA, CISA, and FBI released advisory AA26-251A last week, explicitly naming six AI companies involved in or targeted by malicious distillation campaigns. The advisory urges AI model buyers to implement stricter provenance checks and API abuse monitoring to prevent competitors from extracting proprietary capabilities through "distillation attacks." This marks a shift in how distillation is viewed: no longer just a technique for efficiency, but a vector for intellectual property theft and national security risk. For training researchers, this implies a need for robust defenses against output-based extraction methods.
Inference Costs Drop as Labs Aggressively Cut Prices
Chinese tech outlet 80aj.com reported this week that the "distillation war" and efficiency gains are driving prices down further. Specifically, Zhipu’s new GLM-5.3-Flash model is priced at approximately 0.13 RMB per million tokens (approx. $0.018 USD) under limited-time discounts, undercutting DeepSeek-V4-Flash. This aggressive pricing strategy by Chinese labs is forcing global competitors to optimize their inference stacks rapidly to maintain margin viability, accelerating the adoption of quantization and sparse attention mechanisms in production environments.
Local view
Local stakeholders in China are closely monitoring the "price war" among large language models. Zhihu discussions highlight that GLM-5.3-Flash's new pricing structure makes it one of the most cost-effective options for enterprise developers, particularly for high-volume, low-latency tasks. The consensus among local analysts is that the combination of MoE architectures and distillation techniques has lowered the barrier to entry for high-quality AI services, shifting competition from raw capability to cost-per-token efficiency.
Context & numbers
- Active Parameters: Instella-MoE uses 2.8B active parameters out of 16B total, maintaining high efficiency.
- Training Volume: The Instella model was trained on 7.1T tokens.
- Hardware: Instella-MoE was trained on AMD Instinct MI300X/MI325X GPUs.
- Pricing: GLM-5.3-Flash is listed at ~0.13 RMB/M tokens during promotional periods.
On the radar
- FP8/FP4 Adoption: The recent Megatron-Core technical report highlights Section 5 on Reduced-Precision Training in FP8/FP4 for MoE, suggesting that next-generation efficiency gains will come from lower-precision formats rather than just architectural changes.
- Provenance Standards: Following the AA26-251A advisory, expect increased demand for tools that verify the provenance of training data and model weights to comply with new security guidelines.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.