AI raised these students' homework scores by 18% — then cut their exam scores by 20%
A 30-month study of 26,811 Chinese secondary students pulls apart looking productive from actually learning — and the gap took two years to show up in full.

The finding in one sentence
When secondary-school students started using generative AI for homework, their homework got better and their exam results got worse.
That is the core result of "The Generative AI Learning Penalty," a Centre for Economic Policy Research discussion paper (DP21577) by David Strömberg of Stockholm University and Victor Lei and Yanhui Wu of the University of Hong Kong. It follows 26,811 students in grades 7 to 12 across a single county in central China over 30 months, from September 2022 to June 2025 — a window that brackets the arrival of ChatGPT and the wave of Chinese chatbots that came after it.
This is not a lab experiment or a survey. It is administrative data: real homework marks, real completion times and real closed-book exam results across nine subjects, for tens of thousands of students, month after month. That scale and that length are what make it worth reading — they let the authors watch a slow effect build that shorter studies miss.
The numbers
Once students took up AI tools — Doubao, DeepSeek, Qwen, Ernie Bot and ChatGLM among them, with adoption reaching about 80% — three things moved together:
- Homework scores rose about 18%.
- Homework time fell from 64 minutes to 45 — a drop of roughly 30%.
- Monthly closed-book exam scores fell about 20% within six months.
The exam damage compounded. After about two years, high-stakes entrance-exam scores were down 18% to 24%: a 24% fall in the zhongkao, the high-school entrance exam, and an 18% fall in the gaokao, the national college entrance exam. The full penalty only became visible near the end — which is exactly why the authors argue that earlier, shorter studies looked reassuring.
Who lost the most
The losses were not spread evenly. They were largest in social-science subjects, then STEM, then languages. They were worst for junior students, for high-achieving students, and for boys. And they concentrated in the roughly 80% of AI users whose pattern looked like outsourcing rather than help — very short homework times paired with very high homework marks.
That pattern is the mechanism the paper leans on. The problem is not the tool; it is what the tool lets a student skip. As the authors put it: "For students, completing these tasks efficiently is not the goal; learning from them is." Their conclusion is blunt — generative AI, "likely to become a prevalent technology for education, has a substantial negative impact on student learning."
Why it matters now
The paper has been circulating since June and resurfaced in wider discussion this week. It keeps coming back because it lands squarely in the argument every school, parent and edtech company is having right now: does AI help students learn, or just help them finish? It brings an unusually large and long dataset to that question — 26,811 students over 30 months — and its answer is that finishing and learning came apart.
Two caveats worth keeping in view. First, this is observational data from one county inside one country's exam-driven system; the outsourcing channel is well identified, but a single setting is not the whole world. Second — and the authors make this point themselves — the result is about how the tools were used, not proof that they must be used that way. The same study that measures the penalty also points at the fix: AI that makes the student do the thinking, not AI that does it for them.
Ask Relay — he reads every question himself and replies personally by email.
