Small and On-Device Models: Phi, Gemma, Apple — 2026-09-26
The week's big story is on-device hardware: Qualcomm unveiled its Snapdragon 8 Gen 6 Elite with an upgraded Hexagon NPU built for agentic, on-device AI, while Microsoft quietly killed the "Copilot+ PC" branding even as the AI features remain. On the software side, Hugging Face's Transformers library now supports llama.cpp GGUF quantizations, making 4-bit small models easier to run locally.
Small and On-Device Models: Phi, Gemma, Apple — 2026-09-26
Top developments
Qualcomm unveils agentic-AI smartphone chip at Snapdragon Summit 2026
At its Snapdragon Summit in Hawaii (Sept 23–25), Qualcomm introduced the next-generation flagship Android chip with strengthened on-device AI, framing a shift to "agentic AI" — smartphones running models locally that understand user context and control multiple apps directly, using a personal knowledge graph to unify information across devices. Korean coverage of the "Snapdragon 8 Elite Gen 6" notes the architecture splits roles: the CPU "orchestrates" while the NPU does the "reasoning." This matters directly for small on-device models, since agentic local inference depends heavily on NPU memory capacity and throughput.

Microsoft drops "Copilot+ PC" branding, keeps the AI features
Multiple outlets reported on Sept 25 that Microsoft is quietly retiring the "Copilot+ PC" brand for Windows laptops; a Surface CVP confirmed new Surface machines no longer carry the label but still support all Copilot+ features. The hardware requirements — NPU TOPS thresholds and Windows on-device AI features like Recall-style local processing — remain in place. For the small-model ecosystem, the practical takeaway is that on-device AI differentiation is shifting from marketing badges to actual NPU capability and the models that ship on it.
Hugging Face Transformers now runs llama.cpp GGUF quants
Announced roughly four days ago, the Transformers library now natively supports loading llama.cpp-style GGUF quantized weights. GGUF bundles weights, tokenizer, and chat template into a single file, with multiple quantization levels trading precision for a smaller memory footprint. This lowers the barrier for shipping 4-bit versions of small models (Phi-4-mini, Gemma-class) from a Python stack directly to constrained local hardware.

Running small models directly on phones: Ollama on Android, MLC/MLX on iOS
A hands-on guide published this week walks through running Gemma-family small models on smartphones: Android users install Ollama via Termux and pull a small Gemma model, while iOS users rely on MLC Chat, MLC LLM, Apple MLX or Core ML. 4-bit quantization is highlighted as the key lever to reduce memory footprint enough for phone-grade RAM.
Local view
Korean tech media has been fixated on Qualcomm's agentic-AI pitch from Snapdragon Summit. AI타임스 (AI Times) framed the new premium Android chip as enabling models to run "on the smartphone itself" with context understanding. 조선일보 (Chosun Ilbo) described the "AI phone that knows before you speak," built on personal knowledge graphs. 테크M quoted Qualcomm's framing that mobile AI has evolved from generative to agentic, with the "blossoming imminent". Separately, HelloT reported that Korean robot maker Brils is pushing physical AI with domestically produced NPUs and on-device VLA/robot foundation models.

Context & numbers
- The overall smartphone market is expected to contract 14% in units shipped in 2026, weighed down by a memory shortage — a squeeze that directly affects how much RAM phones can allocate to on-device AI models (CNBC, Sept 22)
- For perspective on quantization economics: Q4_K_M (GGUF) or AWQ 4-bit delivers 1–2% perplexity degradation, 3.5–3.8x speedup, and roughly 4x memory savings versus FP16 (May measurement, pre-cutoff but still the standard reference)
- Microsoft's Copilot+ era números: AI laptops in the category start around $829 (July pricing data, pre-cutoff)
On the radar
- Snapdragon Summit 2026 announcements will trickle into device releases over the coming quarters — watch which phone vendors pair the new Hexagon NPU with small locally-run agentic models
- Snapdragon AI PCs with local processing, battery efficiency and Wi-Fi 7 features are arriving in new devices imminently
- Rumor/flag: Samsung Galaxy AI's on-device stack and how it adapts to the new Qualcomm NPU generation remains a watch item in Korean coverage
- Korean low-power NPU players (e.g., DeepX smart-glass PoC validation) are testing ultra-low-power on-device AI in wearables — a segment to watch
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.