Small and On-Device Models: Phi, Gemma, Apple — 2026-09-10
Apple’s newly announced M6 chip and macOS Golden Gate update are reshaping the on-device AI landscape, with early benchmarks suggesting a significant leap in local inference capabilities compared to Microsoft’s Copilot+ strategy. Meanwhile, South Korean tech firms are aggressively optimizing Vision-Language-Action (VLA) models for domestic NPU hardware to enable real-time humanoid robot control without cloud dependency.
Small and On-Device Models: Phi, Gemma, Apple — 2026-09-10
Top developments
Apple M6 Chip Targets Heavy Local AI Workloads
Apple has officially detailed its next-generation silicon, featuring a 2nm process and a "quad-die" architecture specifically engineered for on-device AI. The new Mac processors are designed to handle cutting-edge large language models locally, leveraging massive unified memory bandwidth to overcome previous bottlenecks. This move directly challenges the current paradigm where heavy AI tasks often default to cloud processing, positioning Apple’s hardware as a primary engine for private, high-performance local inference.

macOS Golden Gate Outperforms Copilot in AI Integration
Recent analysis indicates that Apple’s macOS "Golden Gate" update has surpassed Microsoft’s desktop Copilot in practical utility and seamless integration. Despite Microsoft’s three-year head start with AI-integrated desktops, Golden Gate’s implementation of Siri and AI search tools is considered superior for everyday workflows. This shift suggests that Apple’s long-game approach to embedding Foundation Models directly into the OS architecture is yielding better user experiences than the bolt-on approach seen in some Windows implementations.

Nota Optimizes VLA Models for LG Humanoids on Domestic NPUs
South Korean AI optimization firm Nota has been selected as the lead for the government-backed "K-On-Device AI Semiconductor" project. They are developing real-time optimization techniques for Vision-Language-Action (VLA) models to run on domestic Neural Processing Units (NPUs), specifically targeting LG Electronics' humanoid robots. By eliminating cloud dependency, this initiative aims to achieve sub-millisecond latency for robot control, marking a significant step for local-language AI deployment in industrial robotics.

Acer Unveils Dedicated On-Device AI Hardware at IFA 2026
At IFA 2026 in Berlin, Acer launched a new line of local AI desktops and ultra-light laptops (799g) explicitly marketed for on-device AI computation. The company is pivoting its competitive strategy from raw processing speed to dedicated AI inference capabilities, signaling a broader industry shift where NPU performance becomes a primary differentiator in consumer hardware. These devices are optimized to run small language models locally, reducing latency and privacy concerns for end-users.
Local view
Korean media outlets are closely monitoring the "K-On-Device AI" initiative, highlighting how local firms like Nota are partnering with global giants like LG Electronics and chipmaker Mobilint to create a self-sufficient AI ecosystem. The focus is on reducing reliance on foreign cloud services and leveraging domestic NPU technology to drive real-time applications in robotics and smart devices.
Context & numbers
The push for on-device AI is driven by three core requirements: latency, privacy, and offline availability. Recent deployments of quantized small language models (SLMs) on mobile devices demonstrate that these constraints can be met without significant performance degradation, provided the correct quantization formats (such as GGUF or MLX) are utilized.
On the radar
- Microsoft Copilot+ PC Strategy: Following comments at Build 2026, Microsoft appears to be relaxing restrictive hardware requirements for Copilot+ PCs, potentially allowing older hardware to access newer on-device AI features. This could democratize access to local AI tools previously locked behind specific NPU thresholds.
This content was collected, curated, and summarized entirely by AI — including how and what to gather. It may contain inaccuracies. Crew does not guarantee the accuracy of any information presented here. Always verify facts on your own before acting on them. Crew assumes no legal liability for any consequences arising from reliance on this content.