The AI Trick Your Watch Already Knows
The industry keeps telling us we need dedicated AI chips and next-gen silicon for serious on-device intelligence. That's a lie. A recent demonstration from Better Stack exposed this myth, showcasing a Falcon H1 large language model, a compact 90-million-parameter marvel, running fully offline on an Apple Watch series 6. This isn't some beta test on a developer kit; it's a 2020 device, my series 6, running an LLM locally.
Everything happens on the edge, with no cloud streaming involved. The model streams text at a genuinely brisk 15 tokens per second, proving that even older hardware can deliver a responsive AI experience. Beyond simple text generation, it supports sophisticated tool calls, effortlessly accessing external data like Wikipedia for instant answers, all via the watch’s native voice input. The capital of France? Paris, delivered almost instantly.
This feat becomes even more astonishing when you consider the hardware. The Apple Watch series 6 sports a mere 1GB of RAM and an Apple S6 processor—components widely deemed insufficient for any serious AI workload. Yet, it works. This demonstration isn't just a technical curiosity; it’s a stark reminder that the real bottleneck isn't always hardware, but often our own assumptions about what’s possible.
How 'Impossible' AI Becomes Possible
The magic isn't in some futuristic chip, but in ingenious software. What seems impossible on a 2020 Apple Watch series 6 becomes reality through relentless model compression. Techniques like 4-bit quantization don't just reduce an LLM's memory and storage footprint; they drastically slash it by factors of 4x or even 8x, making massive models compact enough for embedded systems without a catastrophic drop in quality or performance.
This isn't just about shrinking; it's about smart design from the ground up. A new breed of sub-billion parameter models now dominates the edge. We're talking about models like:
- The Falcon H1 (a mere 90 million parameters for the watch demo)
- Gemma 2B
- Phi-4 mini
These architectures are meticulously engineered for efficiency and pair with optimized runtimes and formats such as GGUF, specifically built for robust CPU and low-power inference, bypassing the need for specialized AI accelerators.
Here's the punchline: performant on-device AI on your everyday gadgets isn't solely a brute-force hardware problem. It's the sophisticated interplay of software and model optimization—a relentless pursuit of efficiency that transforms yesterday's devices into tomorrow's AI powerhouses. This combination, not just raw processing power, unlocks truly local, responsive AI experiences.
Big Tech's Deliberate Delay
The purported technical hurdles for widespread on-device AI—battery drain, performance trade-offs—are largely a smokescreen. These are solvable engineering challenges, not insurmountable barriers, yet they conveniently justify Big Tech's glacial pace. The recent demonstration of a 90 million parameter Falcon H1 model running natively on a 2020 Apple Watch series 6, streaming at 15 tokens per second completely offline, proves the capability exists today.
This deliberate delay serves a clear financial purpose: Big Tech’s current business model thrives on cloud-based AI. Cloud operations lock users into lucrative ecosystems, generate invaluable training data, and create predictable subscription revenue streams. On-device AI, by its very nature, directly threatens these established profit centers, offering users true autonomy from perpetual cloud dependence and its associated costs.
Moreover, the industry leverages the illusion of necessity to drive hardware upgrades. Promoting powerful on-device AI as an exclusive, must-have feature for new devices with dedicated NPUs is a strategic sales tactic. This pushes consumers to upgrade their phones and watches, despite older hardware, like the 2020 Apple Watch series 6 running a 90 million parameter Falcon H1 model, already demonstrating robust local AI capabilities.
Enjoying this? Get one like it in your inbox each morning.
one email a day · unsubscribe in two clicks · no third-party tracking
Your Data, Your AI: The Coming Shift
Imagine an AI assistant that responds with zero latency, its knowledge of your life profound yet perfectly secure. This is the imminent reality of true on-device AI. Your personal data, your conversations, your habits—everything remains encrypted and processed locally, never leaving your device for a distant server. No more cloud round-trips, no more data breaches from third parties.
The genie is out of the bottle, and it runs locally. Developers like Better Stack showcasing a 90 million parameter Falcon H1 model operating completely offline on a 2020 Apple Watch series 6 at 15 tokens per second isn't just a tech demo; it's a gauntlet thrown. As consumers witness this undeniable capability, they will demand these privacy-first, offline-capable AI features. Manufacturers will face immense market pressure, forced to adapt their product roadmaps or risk becoming obsolete.
This isn't a future years away, nor a hypothetical. The technology for powerful, personal, on-device AI is demonstrably here, capable of sophisticated tasks on a six-year-old smartwatch. The only remaining barrier is corporate strategy, prioritizing cloud revenue and data harvesting over user autonomy and security. The lie about on-device AI, that it’s inherently limited or impossible, will soon be impossible to maintain.
Frequently Asked Questions
What LLM was used on the Apple Watch in the demo?
The demo featured Falcon H1, a highly efficient 90 million parameter model. Its small size and hybrid architecture are key to its performance on constrained hardware.
What are the main benefits of on-device AI?
The primary benefits are enhanced data privacy (data never leaves your device), lower latency for faster responses, offline functionality without a network connection, and reduced costs by avoiding cloud processing fees.
Why are device makers hesitant to widely adopt on-edge LLMs?
Manufacturers face challenges like battery consumption and hardware limitations. They also have business incentives to favor cloud-based AI, which keeps users in their ecosystem and creates opportunities to sell new hardware with dedicated AI chips.
Does this mean my current phone can run a local LLM?
Yes, it's increasingly possible. With optimized models like Llama 3.2 and runtimes like llama.cpp, many modern smartphones and even older devices can run capable LLMs locally, though performance will vary.

