The Homework Got Better. The Students Got Worse.
A study of 26,000 students found AI homework help raised homework scores 18% — and cut exam scores 20%. It's the clearest example yet of today's running theme: a number rising while the thing it measures falls.

We spent today on a single idea: that a number can go up while the thing it is meant to measure goes down. In AI, that shows up as benchmark scores that mislead. It turns out the cleanest example of the same trap isn't in a lab. It's in a classroom.
A study of roughly 26,000 Chinese secondary students, whose findings drew a fresh wave of attention this week, tracked what happened when they started using generative AI to do their homework. The homework got better. The students got worse. The paper's title does not hedge: "The Generative AI Learning Penalty," by David Strömberg, Victor Lei and Yanhui Wu of Stockholm University and the University of Hong Kong.
The two numbers
The headline is a split. Homework scores among AI users rose about 18%, and the time spent on each assignment fell by roughly 30% — from a reported 64 minutes to 45. By the metric a teacher sees first, these were improving students working more efficiently.
Then the researchers looked at closed-book exams, where no AI is available. Within six months, AI users' monthly exam scores had dropped about 20% relative to classmates who didn't use the tools. Measured across a full two-year window, entrance-exam results fell somewhere between 18% and 24%. The homework number and the exam number moved in opposite directions, and the exam is the one that was actually testing whether anyone had learned anything.
What the gap is made of
The mechanism the researchers point to is blunt, and they give it a name: homework outsourcing. Around 80% of the AI users showed its signature — finishing assignments unusually fast while still scoring highly. The work was getting done; the learning the work exists to produce was being skipped. A homework score was supposed to be a proxy for understanding, and once a tool could generate the proxy directly, the two came apart.
The damage wasn't evenly spread. The study reports the steepest losses in social-science subjects, then STEM, then languages, and found the effect most pronounced among younger students, high achievers, and boys — with the high-achiever finding the most uncomfortable, because those are exactly the students a good homework score is supposed to identify.
The caveats, honestly
This is one study, drawn from one country's exam-heavy education system, and it is not a randomised trial. It is, though, more than a simple correlation: the researchers used a difference-in-differences design that exploits the staggered timing of AI adoption across students — a method built to get at cause rather than coincidence — and they state their result in causal terms. What it still cannot do is prove that any one student would have scored higher without the tool, only that the pattern across nearly 27,000 of them is stark and consistent. And "AI use" here means outsourcing the work, not using AI to explain a concept you then practise — the study is about substitution, not assistance.
Those caveats narrow the claim. They don't dissolve it, because the core finding is not really about AI at all. It is about what happens when you optimise for a measurement instead of the thing the measurement stands for — which is the same failure that lets a language model top a leaderboard it has memorised, and the same reason a rising number is worth less than the question of how it was produced. The students got very good at the homework. That was the problem.
Ask Relay — he reads every question himself and replies personally by email.
