The phone learns to think for itself
At Qualcomm’s Snapdragon Summit in Maui on October 24, 2023, an engineer ran a seven-billion-parameter language model on a smartphone chip with no network connection. The model — Meta’s Llama 2 — generated twenty tokens per second. Not from a server farm, not via satellite link, but from the chip in the device itself. The phone, in its 147th year, was thinking.
The chip was the Snapdragon 8 Gen 3, Qualcomm’s latest 4-nanometer flagship, and AI was its organizing principle. Previous generations had included neural processing units for narrow tasks — face recognition, photo enhancement. The 8 Gen 3’s Hexagon AI Engine was built for something different: running large language models of three to thirteen billion parameters locally, no cloud call required (Engadget). It could also generate a Stable Diffusion image in under a second — a benchmark that had required a desktop GPU twelve months earlier.
Google had placed a parallel bet. The Pixel 8 Pro launched in October 2023 with the Tensor G3 chip, and on December 6, 2023, Google pushed Gemini Nano to it — its smallest model, purpose-built for on-device inference, requiring no network call. The initial features were modest: Summarize in Recorder, which could condense a recorded conversation without touching the internet, and Smart Reply suggestions in Gboard (Google Blog). Modest, but foundational. The AI was in the device — not in a data center in Oregon.
Samsung arrived in San Jose on January 17, 2024, and made the same capability a billboard. The Galaxy S24 series, the second Android device family to carry Gemini Nano, launched under the slogan “Galaxy AI.” Its Magic Compose feature could rewrite a text message in styles ranging from “excited” to “lyrical” to “formal,” entirely on the phone, with nothing landing on a third-party server (Samsung Newsroom). The pitch was partly capability, partly something older: privacy. Your messages, rewritten by a language model, never left your pocket.
This was what separated the new class from what had come before. Siri had been on iPhones since 2011; Google Assistant since 2016. Both worked by routing audio to cloud servers and returning answers — fast when it worked, and wholly dependent on a connection. On-device LLMs cut that wire. A 7B model at 20 tokens per second is slower than a good cloud call, but it works in a tunnel, in airplane mode, in places where you’d rather not route your queries through someone else’s infrastructure.
What the 2023 shift announced, architecturally, was a change in where intelligence lives. For three generations of smartphone AI, the device was a terminal: it captured your words, sent them away, displayed what came back. Pulling the computation onto the device made the phone not a window into intelligence but a seat of it.
The telephone Bell patented in 1876 was a conduit — a wire between two voices. One hundred and forty-seven years later, it had become something he wouldn’t have recognized: a machine that interprets the world on its own terms, without asking anyone for permission.
Sources
- Engadget: Qualcomm’s Snapdragon 8 Gen 3 brings on-device generative AI — Snapdragon 8 Gen 3 announcement October 24, 2023; Llama 2 running at 20 tokens/second on-device; Hexagon AI Engine for 3B–13B parameter models.
- Google Blog: Pixel 8 Pro December 2023 Feature Drop — December 6, 2023 Gemini Nano rollout; Pixel 8 Pro first smartphone engineered for Gemini Nano; Tensor G3 chip; Summarize in Recorder and Smart Reply features.
- Samsung Global Newsroom: Samsung and Google Cloud generative AI for Galaxy S24 — January 17, 2024 Galaxy S24 launch; Galaxy AI branding; Magic Compose on-device text rewriting; privacy-first framing.