Why AI Keeps Getting Cheaper: The Economics Behind the Price Cuts
Another week, another price cut — flagship AI costs a fraction of what it did in 2023. Four forces are driving the collapse, and one paradox explains why your bill might still be going up.

If you follow AI at all, you have seen the pattern: another week, another price cut. This week, OpenAI dropped the price of its flagship model by more than a fifth — the latest move in a steady run of cuts across the whole industry. Rewind two years and the same intelligence cost many times more. Almost nothing else in technology is getting cheaper this fast.
That is not a fluke or a loss-leader stunt. It is the product of four forces pushing in the same direction at once. Here is what is actually driving the collapse in the price of AI — and the catch that means your bill may not be falling at all.
First, the numbers
AI is priced per token — very roughly, a token is a chunk of text about three-quarters of a word. Providers charge separately for the tokens you send in (the prompt) and the tokens the model sends back (the answer), and the output side is usually the pricier of the two.
The direction of travel is steep and consistent. Flagship models that cost tens of dollars per million tokens at launch in 2023 now cost a few dollars, and the cheap-and-fast tier has fallen further still. The specific cuts change most months — some, like OpenAI's latest, are explicitly time-limited promotions — but the direction of travel does not.
Why it keeps happening
Competition, with a floor near zero. Every major lab is racing to be the default, and price is the bluntest lever they have. What makes this fiercer than an ordinary price war is the presence of capable open-weight models — systems anyone can download and run themselves. When a free-to-run model is good enough for a task, no provider can charge much above the cost of the electricity to serve it. Open weights set a floor, and the floor keeps dropping.
Hardware gets more per watt. Each generation of AI accelerator does more work for the same power and the same rack space. The chips are not cheap, but the amount of useful output squeezed from each one keeps rising, and that shows up directly in the per-token price.
The models themselves got leaner. A great deal of 2026's research went into making smaller models behave like larger ones — through distillation, better training data, and architectures that only switch on the parts of the network a given question needs. A model that is a fraction of the size but nearly as capable is a fraction of the cost to run.
Serving got smarter. Behind the scenes, providers have become far better at the plumbing: batching many users' requests together, caching work that repeats, and running models at lower numerical precision without meaningfully hurting quality. None of this changes what the model can do; all of it lowers the cost of doing it.
The catch: cheaper tokens, bigger bills
Here is the part that surprises people. The falling price per token has not, for many companies, produced a falling bill. It has often produced the opposite.
The reason is an old one, and it has a name: the Jevons paradox. When something useful gets cheaper, people use dramatically more of it. Cheap tokens made it sensible to point AI at jobs that were never worth it before — reading every document, drafting every reply, running an agent that thinks in long loops rather than answering in one shot. Those agentic workloads, in particular, can burn through tokens at a rate that a single chat answer never approached.
So two things are true at once: each unit of AI is getting cheaper, and total spending on AI is climbing. A lower sticker price is an invitation to use more, and most of the market has accepted the invitation.
What it means, and what to watch
For anyone building on AI, the practical lesson is that model choice is now a real budget decision, not a rounding error. The gap between the flagship tier and the cheap-and-fast tier is often wide enough that using the right-sized model for each job — rather than the best model for every job — is one of the biggest levers on cost.
Three things are worth watching if you want to know where prices go next:
- The open-weight frontier. The better free-to-run models get, the harder it is for anyone to charge for the same capability. Watch how close the best open model sits to the best closed one.
- The output-token gap. Output is where the money is. Cuts to output pricing matter more to a real bill than the headline input number that usually leads the announcement.
- Whether efficiency outruns usage. Prices per token will keep falling. Whether your total bill follows depends entirely on whether your appetite grows slower than the price drops — and for most, so far, it has not.
Cheap intelligence is the defining commercial fact of this AI cycle. Just don't mistake a falling price for a falling cost.
Ask Relay — he reads every question himself and replies personally by email.
