CrewCrew
FeedSignalsMy Subscriptions
Get Started
Open Weights Outside China: Llama, Mistral, Gemma

Open Weights Outside China: Llama, Mistral, Gemma — 2026-09-02

  1. Signals
  2. /
  3. Open Weights Outside China: Llama, Mistral, Gemma

Open Weights Outside China: Llama, Mistral, Gemma — 2026-09-02

Open Weights Outside China: Llama, Mistral, Gemma|September 2, 2026(3h ago)4 min read8.5AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

The open-weight landscape outside China saw significant activity this week, headlined by Z.ai’s release of GLM-5.3 weights and Mistral AI’s strategic pivot to hosting Chinese models. While Western labs like Meta and Google maintain steady ecosystems, the "capability premium" has collapsed as Chinese open-weight models achieve frontier-level performance at lower costs. Local inference communities are increasingly adopting these high-parameter models despite licensing complexities.

Open Weights Outside China: Llama, Mistral, Gemma — 2026-09-02


Top developments

Source image
Source image

local-ai-zone.github.io

local-ai-zone.github.io


Z.ai releases GLM-5.3 open weights, challenging Western dominance

On August 29, Z.ai (formerly Tsinghua KEG) released the weights for GLM-5.3 and GLM-5.3-Flash on Hugging Face. The flagship GLM-5.3 is a 753B parameter MoE model, while the Flash variant is a 320B multimodal model released under an MIT license. This release is notable for its scale—1.4TB of weights—and its competitive benchmarks, with Terminal Bench scores jumping from 4.6 to 28.3 compared to its predecessor. The move reinforces the trend of Chinese labs offering frontier-class open weights that undercut the cost-per-token of Western counterparts like Llama and Gemma.

Source image
Source image

blogs.nvidia.com

blogs.nvidia.com


Mistral AI hosts Chinese models, shifting focus to sovereignty

In a move that sparked debate within the European AI community, Mistral AI began hosting GLM-5.2, a model from Chinese lab Z.ai, within its API and Vibe platform. French media outlets like Frandroid reported that while some community members viewed this as a betrayal of European sovereignty principles, others praised the pragmatic approach to offering best-in-class performance regardless of origin. This aligns with Mistral’s broader repositioning from a direct competitor to OpenAI to a facilitator of European digital sovereignty, where the platform itself becomes the sovereign layer rather than the underlying model weights alone.


Five open-weight releases in nine days collapse the price floor

Between August 20 and 28, five major labs including Z.ai, Alibaba, Tencent, MiniMax, and DeepSeek released open-weight models with 1M token context windows at "Flash tier" prices. Requesty.ai data indicates that model launch chatter doubled week-over-week during this period. The simultaneous release of GLM-5.3-Flash, Qwen3.8-Flash, Hy4, and others has effectively eliminated the "capability premium" previously enjoyed by Western open-weight models like Llama 4 and Gemma 4. For developers, this means access to near-frontier intelligence without the associated cost or latency penalties of closed-source APIs.


Hugging Face reports massive shift in GGUF download trends

Hugging Face’s "State of Open Models" report highlights a dramatic shift in local inference preferences. Qwen models now account for 39.6 million GGUF downloads per month, nearly double that of Google’s Gemma (20.8 million) and more than five times that of Meta’s Llama (7.5 million). Despite Llama-derived repositories slightly outnumbering Qwen’s on the platform, the download velocity favors Qwen, indicating stronger developer adoption for local deployment. This trend underscores the erosion of Llama’s traditional dominance in the open-weight ecosystem as non-US models offer superior performance-to-size ratios.


Local view

France: Sovereignty vs. Pragmatism French tech outlet Developpez analyzes Mistral AI’s strategic evolution, noting that the company has moved away from competing solely on model performance to positioning itself as the guardian of European digital sovereignty. By hosting diverse models, including Chinese ones, Mistral prioritizes utility and compliance over ideological purity in model sourcing. Meanwhile, Frandroid captures the community’s divided reaction, highlighting the tension between supporting local innovation and accepting the superior performance of imported open weights.

Japan: Focus on Specialized Open Models Japanese media continues to highlight domestic efforts alongside global trends. AI Revolution provided a detailed breakdown of NII’s LLM-jp-4 33B, emphasizing its Apache 2.0 license and Japanese-specific performance benchmarks that surpass gpt-oss-20b. Additionally, GIGAZINE covered the GLM-5.3 release with significant interest, noting its potential impact on local developers who can now access high-performance models under permissive licenses. These reports reflect a growing sentiment that "purely domestic" AI is less about proprietary weights and more about localized fine-tuning and regulatory alignment.


Context & numbers

  • Download Dominance: Qwen leads GGUF downloads at 39.6M/month, followed by Gemma at 20.8M/month, and Llama at 7.5M/month.
  • Model Sizes: Recent releases include GLM-5.3 (753B parameters), GLM-5.3-Flash (320B parameters), and Kimi-K3 (~2.8T parameters).
  • Licensing: Mistral’s open models remain primarily Apache 2.0. Z.ai’s GLM-5.3-Flash is MIT licensed, while the flagship GLM-5.3 has specific usage terms detailed in its model card.
  • Performance: GLM-5.3 shows a Terminal Bench score improvement from 4.6 to 28.3 over its predecessor, signaling rapid iteration in agentic capabilities.

On the radar

  • Debian LLM Vote Results: The Debian community recently concluded a vote on LLM usage in project documentation, with "neither endorsed nor prohibited" winning. This signals a tentative acceptance of AI tools in open-source governance, potentially influencing how open-weight models are integrated into development pipelines.
  • NVIDIA Nemotron Updates: NVIDIA continues to promote its Nemotron 3.5 Lightning model for local agentic AI, emphasizing its free commercial use and modification rights. Expect increased integration of Nemotron into enterprise-grade local stacks as NVIDIA pushes its hardware-software synergy.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow do Western labs plan to respond to price drops?
  • QWhat hardware is needed to run GLM-5.3 locally?
  • QWill Mistral face regulatory pushback for hosting?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.