In AI Inference, the Bottleneck Is Moving Data Out of Memory — Here's Why That Changes the Hardware
Training built the AI boom. Running the models is now what drives it, and a detailed IEEE Spectrum report shows that shift is rewriting chip design. For much of the work, the slow part is moving data out of memory, not doing sums, which is why Nvidia, Amazon and a crop of startups are all attacking memory.

The takeaway: For years the AI hardware race was about training, building ever-bigger models. A long IEEE Spectrum feature by Matthew S. Smith argues that in 2026 inference, actually running those models for users and agents, has come to the forefront, though labs are still training ever-larger models too. Inference stresses different parts of a chip. For the stage where a model writes its answer, a major bottleneck is how fast data can move out of memory, and that helps explain several of the unexpected hardware alliances of the past year.
Why inference suddenly matters so much
Three things are piling up, according to the report:
- People use the models. They became useful, so demand grew.
- Reasoning models think at length. A model with high reasoning effort "can produce up to 20 times as much text" as one with low or no effort, and all of that text has to be generated.
- Agents never clock off. Agentic AI runs inference around the clock, not just when someone types a question.
Matt Kimball, principal data-center analyst at Moor Insights & Strategy, put it bluntly to IEEE Spectrum: "It's like training is yesterday's news."
The two halves of answering a prompt
Every reply from a large language model happens in two phases, and they want different hardware:
- Prefill is the model reading your prompt. It handles every token at once, which splits neatly across many parallel cores. That's why GPUs, built for parallel graphics work, became the default AI chip.
- Decode is the model writing its answer, one token at a time. For each new token it has to read the model's weights plus a growing scratchpad of the conversation so far (the "KV cache"), which the report says can swell to dozens of gigabytes.
Decode is the problem. The chip spends much of its time waiting for data to arrive from memory. IEEE Spectrum cites research finding that Nvidia H100 GPUs running open-source LLMs "sit idle 50 to 80 percent of the time."
Everyone is attacking memory, in different ways
- Stack compute on memory (d-Matrix). Its Raptor accelerator sits directly on top of the memory chips. Founder and CTO Sudeep Bhoja says that cuts the distance data travels to "micrometers instead of millimeters."
- Stretch the memory link (Majestic Labs). High-bandwidth memory can only work up to 2 or 3 millimetres from the processor, says Majestic cofounder Shahriar Rabii. Majestic claims its interface can carry data about a metre, letting one server rack reach up to 128 terabytes of cheaper standard memory, against about 20 TB of HBM in Nvidia's GB300 NVL72 rack. Memory analyst Jim Handy tells the magazine high-bandwidth memory costs two to three times as much as standard DRAM.
- Make memory faster (SK Hynix and Samsung). HBM4 is now in production and will be used in Nvidia's Vera Rubin GPU, due in the second half of 2026. SK Hynix's Hoshik Kim says HBM4 "will decisively break the memory bottlenecks constraining AI inference today."
- Put the memory on the chip (Nvidia, Cerebras). Nvidia bought intellectual property and hired talent from Groq at the end of 2025. Its Groq 3 language-processing unit trades raw compute for 500 megabytes of on-chip SRAM. Nvidia's Ian Buck says it "has seven times the memory bandwidth of the GPU." Cerebras goes further: its wafer-sized WSE-3 chip holds 44 gigabytes of SRAM.
The big players are splitting the job across two chips
Nvidia and Amazon Web Services seem to agree: use one kind of chip for prefill and another for decode.
- Nvidia plans to run prefill on its Rubin GPUs and hand decode to the Groq 3 LPU.
- AWS plans to pair its own Trainium chips for prefill with Cerebras's wafer-scale engine for decode.
Buck's summary to IEEE Spectrum: "To do modern AI inference, you need all the chips."
Doing more with fewer bits
Software is changing too. Quantization stores a model's numbers at lower precision, so the same memory holds more model. Nvidia says that when it converted DeepSeek-R1 to its new 4-bit NVFP4 format, scores on seven major benchmarks dropped by less than one percent while performance improved threefold. AMD, Intel and Qualcomm back a rival 4-bit format, MXFP4, which Nvidia also helped develop.
Some startups are going further:
- Tensordyne uses a logarithmic number system so its chip can add where it would otherwise multiply. It says its Napier hardware can reach up to 1,300 tokens per second per user on less than a tenth of the power of comparable Nvidia hardware. Tensordyne expects its first hardware in 2027.
- Etched hard-wires the transformer architecture into silicon. It claims its Sohu accelerator runs Meta's Llama 70B at 500,000 tokens per second, but the chip can't run models that move away from the standard transformer. Etched shipped its first rack in August.
Treat the startup figures as company claims. The startup hardware is not yet widely deployed.
Why it matters
In our view, this report is a useful corrective to a chip story told mostly in GPU counts. For how quickly AI can write its answers, memory and data movement are a major constraint, not just raw compute, and that context helps explain why Nvidia paid for Groq's technology and people, and why Amazon turned to Cerebras despite having its own chips. As Kimball says of agents: "These things work 24 hours a day; they don't go home at five at night like we do." If that demand holds, inference hardware is likely to diversify rather than settle on a single winner.
Ask Relay — he reads every question himself and replies personally by email.
