Apple's AI strategy is silicon: the new Mac Studio can run huge models on your desk
The M6 goes 2nm and the M5 Ultra ships with up to 512GB of unified memory — enough, Apple says, to run models with hundreds of billions of parameters entirely on-device. While the frontier race happens in the cloud, Apple is betting on AI that never leaves your machine.
The same week it was reported to be cutting jobs to refocus on AI, Apple has shown where that focus actually points — and it is not a chatbot. It is silicon built to run AI models on your own desk.
Apple introduced two chips today: the M6, its first processor built on a 2-nanometer process, and the M5 Ultra, which it calls its most powerful chip ever. Both are pitched, explicitly, at running AI locally.
The M6
The M6 lands first in a new Mac mini. It pairs a 12-core CPU with a 12-core GPU that now has a Neural Accelerator in every core, plus a dual 16-core Neural Engine Apple says delivers up to twice the peak compute of previous generations. The company claims roughly 30% more peak GPU compute for AI than the M5, with up to 32GB of unified memory at 170GB/s. Apple's own framing is the tell: it is built to run "LLMs on device for secure and private agentic tasks."
The M5 Ultra
The headline part is the M5 Ultra, in a new Mac Studio. It is Apple's first quad-die design — two dual-die M5 Max chips fused with its UltraFusion interconnect — scaling to a 36-core CPU, an 80-core GPU, and a 32-core Neural Engine. Apple claims up to 4.5x the peak GPU compute for AI of the M3 Ultra. But the number that matters most for AI is the memory: up to 512GB of unified memory at 1.2TB/s. That is enough, Apple says, to "run huge LLMs with hundreds of billions of parameters entirely on device."
"M5 Ultra features a massive GPU, now with Neural Accelerators, and more unified memory bandwidth, pushing the boundaries of what a desktop can do." — Sri Santhanam, Apple VP of Silicon Engineering Group
The strategy hiding in the spec sheet
512GB of unified memory is not a number you need for a faster spreadsheet. It is a number you need to hold a very large language model in memory and run it locally. While OpenAI, Anthropic and Google race to build the best model in the cloud, Apple is quietly building the best consumer hardware to run a model off the cloud — on a machine you own, with data that never leaves it.
That reframes this week's other Apple story. Apple looks like the laggard in the AI race only if you assume the race is about chatbots. Its actual bet is that a lot of valuable AI will run on-device — private by construction, with no per-token bill and no dependency on a supplier. It is the individual's version of the sovereign AI that governments and companies are suddenly paying for: keep the model, the data and the compute on your own hardware.
The caveats are real. The performance figures are Apple's own; the 512GB configuration that runs the biggest models is a top-of-the-range Mac Studio, not a cheap laptop; and an on-device model still will not match a frontier cloud model for raw capability. But the direction is a deliberate strategic choice, not a catch-up. The rest of the industry is selling access to intelligence. Apple is selling the machine that keeps it in the room with you.
Ask Relay — he reads every question himself and replies personally by email.
