CrewCrew
FeedSignalsMy Subscriptions
Get Started
Small and On-Device Models: Phi, Gemma, Apple

Small and On-Device Models: Phi, Gemma, Apple — 2026-09-24

  1. Signals
  2. /
  3. Small and On-Device Models: Phi, Gemma, Apple

Small and On-Device Models: Phi, Gemma, Apple — 2026-09-24

Small and On-Device Models: Phi, Gemma, Apple|September 24, 2026(2h ago)3 min read8.7AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

The biggest ambiguity-clarifier this week: Qualcomm says smartphones will run 30B-parameter AI models thanks to memory redesign, Microsoft quietly killed the Copilot+ PC branding while keeping NPU requirements, and Hugging Face made GGUF llama.cpp quants runnable directly in Transformers. Korean media report a surge in on-device AI stocks and note Samsung validated LPDDR6 on Qualcomm APs.

Small and On-Device Models: Phi, Gemma, Apple — 2026-09-24


Top developments


Qualcomm: smartphones can run 30B-parameter on-device models

At Snapdragon Summit 2026, Qualcomm EVP Chris Patrick said on-device AI models — previously constrained by smartphone memory capacity — can grow to 30 billion parameters through data placement design, since more capable AI agents need larger models. Allocating model data across storage and memory could push on-device far beyond today's 3–7B-class small models that dominate phones. Qualcomm simultaneously unveiled a next-generation premium Android chip focused on agentic AI that runs models on the device while understanding user context.

Qualcomm next-gen on-device AI chip announcement
Qualcomm next-gen on-device AI chip announcement


Microsoft quietly retires the Copilot+ PC brand

Multiple outlets confirmed today that Microsoft is dropping "Copilot+ PC" branding from its new 2026 Surface devices, though the underlying AI features and hardware requirements remain intact — Windows 11's 40 TOPS NPU, 16 GB RAM and 256 GB storage thresholds are unchanged. It matters for on-device AI because the NPU spec bar that standardized local model inference across OEM laptops survives even as the marketing label disappears.

Microsoft Surface laptops dropping Copilot+ branding
Microsoft Surface laptops dropping Copilot+ branding


Hugging Face: Transformers now runs llama.cpp GGUF quants

Hugging Face announced that Transformers can now load llama.cpp quantized GGUF files, which package weights, tokenizer metadata and chat template in a single file, with flexible quantization levels trading precision for a smaller memory footprint. This is a practical win for small-model deployment: engines like Phi-4-mini and Gemma in 4-bit can flow between the two most popular local inference stacks.

Hugging Face blog on Transformers running llama.cpp quantizations
Hugging Face blog on Transformers running llama.cpp quantizations


Running Gemma and Ollama directly on phones

A fresh 2026 tutorial walks through running small Gemma models on Android via Termux + Ollama, and on iOS via MLC Chat, MLX or Core ML apps, noting that 4-bit quantization cuts the memory footprint enough to make phone-side inference practical. It confirms the current status quo: on phones, on-device workloads stay in the 1B–4B class rather than pushing frontiers.

Local LLM comparison on smartphones guide
Local LLM comparison on smartphones guide


Local view

Korean media are heavily focused on the hardware side of on-device AI. Pinpoint News reports that on-device AI concept stocks (Gaonchips, Daeduck Electronics) surged on Sept 22–23 on expectations that edge compute demand would spread from smartphones to XR, robots and home appliances. Seoul Economic Daily reports Samsung validated its LPDDR6 mobile DRAM — a world-first — on Qualcomm's latest mobile AP, framed explicitly as a bid to lead the on-device AI market, which ties directly to running larger quantized models on phones. HelloT covers Korean robotics firm Brills betting on a domestic NPU-based on-device physical AI stack with vision-language-action and robot foundation models.


Context & numbers

  • 30B: the parameter scale Qualcomm claims next-gen phones can run, enabled by memory data-placement design — a leap from today's typical 3–7B on-device models.
  • 40 TOPS NPU / 16 GB RAM / 256 GB storage: Windows on-device AI hardware requirements that survive the Copilot+ branding retreat.
  • ~9%: Samsung Exynos mobile AP global market share, its highest since 2024, boosted by Galaxy S26 adoption — more Samsung-branded silicon available for on-device AI.
  • 4-bit quantization remains the standard technique for cutting model memory footprint on phones and laptops.

On the radar

  • Samsung's LPDDR6-with-Snapdragon pairing suggests next Galaxy and Snapdragon devices will advertise larger on-device model runtimes; watch for parameter-count claims at launch.
  • How OEMs rebadge AI-laptop features without "Copilot+" naming — Windows Intel and AMD vendors will likely follow Surface's lead.
  • Google's new $899+ "Googlebook" laptops lean on AI PC positioning and even showcase Microsoft Copilot — watch whether Gemini/Nano models run locally on them.
  • Rumor-level signal: Korean on-device AI equities rallying hard suggests incoming Korean NPU-forward device announcements; treat market chatter, not confirmed product dates.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QHow does Qualcomm's 30B model affect battery life?
  • QWhy did Microsoft drop the Copilot+ PC brand?
  • QWhat are the performance limits of 1B-4B models?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.