CrewCrew
FeedSignalsMy Subscriptions
Get Started
Small and On-Device Models: Phi, Gemma, Apple

Small and On-Device Models: Phi, Gemma, Apple — 2026-09-08

  1. Signals
  2. /
  3. Small and On-Device Models: Phi, Gemma, Apple

Small and On-Device Models: Phi, Gemma, Apple — 2026-09-08

Small and On-Device Models: Phi, Gemma, Apple|September 8, 2026(1h ago)3 min read8.5AI quality score — automatically evaluated based on accuracy, depth, and source quality
0 subscribers

Recent developments highlight the rapid maturation of on-device AI, with a new comparative analysis identifying the top 5 open-source Small Language Models (SLMs) for edge devices, including Llama 3 8B, Phi-3 Mini, and Gemma 2. In South Korea, significant progress is being made in hardware-software co-optimization, with companies like Nota securing government-backed projects to optimize Vision-Language-Action (VLA) models for LG Electronics' humanoids using domestic NPUs. Meanwhile, industry reports emphasize that on-device deployment continues to offer substantial cost reductions compared to cloud-based inference for high-volume workloads.

Small and On-Device Models: Phi, Gemma, Apple — 2026-09-08


Top developments


Press Farm Ranks Top Open-Source SLMs for Edge Devices

A recent analysis published on September 8, 2026, identifies the top five open-source Small Language Models (SLMs) suitable for edge devices. The report compares Llama 3 8B, Microsoft's Phi-3 Mini, Google's Gemma 2, Mistral NeMo, and Qwen 2.5. The study focuses on quantization benchmarks specifically tailored for mobile and IoT applications, providing critical data for developers choosing models for resource-constrained environments. This underscores the ongoing competition among major AI labs to provide efficient, quantized versions of their flagship small models for on-device use.

Chart comparing top 5 open-source small language models for edge devices
Chart comparing top 5 open-source small language models for edge devices

press.farm

press.farm


Nota Optimizes VLA Models for LG Humanoids with Domestic NPU

In a significant move for the Korean on-device AI ecosystem, AI optimization company Nota has been selected as the lead agency for a government "K-On-Device AI Semiconductor" national project. As reported on September 7, 2026, Nota will collaborate with LG Electronics and Movilint to optimize Vision-Language-Action (VLA) models for LG's humanoid robots. The initiative aims to run these complex models directly on domestic Neural Processing Units (NPUs), reducing reliance on cloud connectivity to improve response speeds and security. This represents a concrete application of small model optimization in robotics, moving beyond text-only LLMs to multimodal physical agents.

Nota's AI optimization technology applied to LG Electronics' humanoid robots
Nota's AI optimization technology applied to LG Electronics' humanoid robots


NPU Efficiency Gains Highlighted in Korean Tech Media

Korean tech outlets are increasingly focusing on the energy efficiency of on-device NPUs compared to traditional GPUs. A report from September 6, 2026, notes that on-device NPUs can reduce power consumption by up to 70% compared to GPU-centric approaches for AI inference. As the industry shifts from large-scale training to real-time inference at the service site, this efficiency is becoming a key competitive metric. This trend supports the viability of running larger SLMs like Gemma or Phi variants on mobile devices without excessive thermal throttling or battery drain.

Graph illustrating power consumption reduction of on-device NPUs versus GPUs
Graph illustrating power consumption reduction of on-device NPUs versus GPUs


Local view

South Korean media is closely watching the integration of domestic NPU technology with advanced AI model optimization. The partnership between Nota, LG Electronics, and Movilint is being covered extensively by outlets such as Etoday and TechWorld. These reports highlight the strategic importance of reducing cloud dependency for real-time AI tasks, particularly in robotics and autonomous systems. The narrative emphasizes that local stakeholders are prioritizing "K-NPU" compatibility to ensure that optimized small models can run efficiently on homegrown hardware, fostering a self-sufficient on-device AI supply chain.


Context & numbers

  • Cost Efficiency: On-device deployment continues to deliver 70-90% cost reduction for high-volume, predictable workloads compared to cloud APIs, even as API prices have dropped 40-70% across major providers in 2026.
  • Benchmark Highlights: Recent benchmarks show Phi-4-mini (3.8B parameters) achieving an 83.7% score on ARC-C, the highest in its size class. Gemma 3 4B posts an 89.2% score on GSM8K math reasoning.
  • Power Savings: On-device NPUs offer up to 70% power reduction compared to GPU solutions for equivalent AI inference tasks.

On the radar

  • Windows 11 AI Features: Users are testing which Copilot+ features actually run on desktop PCs with varying GPU capabilities, with reports suggesting some features still require specific high-end hardware despite marketing claims.
  • Google System Services Update: The September 2026 update includes AI-guided onboarding and improvements to laptop file access, potentially expanding the utility of on-device AI services on Android and ChromeOS devices.

This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.

Explore related topics
  • QWhich SLM performed best in the edge benchmarks?
  • QHow do domestic NPUs compare in speed to GPUs?
  • QWhen will LG's humanoid robots launch publicly?
  • QWhat quantization methods were used in the test?

Powered by

CrewCrew

Sources

Want your own AI intelligence feed?

Create custom signals on any topic. AI curates and delivers 24/7.