That Viral AI-Tutor Number Doesn't Hold Up — What the Dartmouth Study Actually Found

The takeaway: A study climbing the developer forums claims an "AI tutor" produced a 0.71–1.30 standard-deviation exam improvement in a Dartmouth statistics course. That number is a model extrapolation from an observational pilot, written by the undergraduate who built the tool — and the paper itself says so. But buried under the headline is a genuinely interesting result: 90% of students voluntarily used an optional, ungraded study tool — and the part that tracked exam gains wasn't the chatbot. It was written-answer questions marked by an LLM.
The claim, and what it actually is
The paper — "Balancing Efficacy and Engagement in Interactive Texts" — was presented on 28 June at the Intelligent Textbooks workshop in Seoul, and hit the Hacker News front page this weekend under a headline touting the 0.71–1.30 SD effect size.
The details matter. The sole author, Jonah Bard, is a Dartmouth undergraduate who built the tool being studied (a web textbook called Phosphor) and submitted the story to Hacker News himself. The venue is a workshop short paper — reviewed, but not a peer-reviewed journal. The study is not a randomised trial: it's an observational analysis of a pilot in an introductory statistics course (151 students enrolled, 143 by term's end; spring term), where the tool was an optional, ungraded alternative to assigned readings.
And the headline number isn't a measured effect. It's a statistical extrapolation — the predicted exam gap between a hypothetical student who completed every lesson and review, and one who did nothing. The paper's own labels for the two endpoints: the 1.30 figure is "selection-inflated", the 0.71 "over-adjusted". The actually-measured contrasts are far more modest: tool users versus non-users scored a statistically non-significant d=0.36 on the final; the only comparison that survived multiple-testing correction (d=0.66) came from the group the paper itself calls "the most self-selected in the study". Hacker News's statisticians were blunter — "calling this an 'effect size' is just nonsense", ran the top skeptical thread, noting the full-engagement endpoint describes roughly 16 students.
To the author's credit, the paper hedges honestly — "self-selection is the central threat" is his sentence, not ours — and he engaged with the critics in the thread. His autumn plan, though, is a grade-attached deployment — still not randomised; a proper randomised trial is something he says he'd "love to run at some point".
The result that does hold up
Strip out the disputed number and two findings look real and interesting:
Students actually used it. 90.2% voluntarily engaged with an optional, ungraded tool, against the 10–15% reading-compliance baseline reported for this same course. Whatever the exam effect, that adoption gap is the practical story for anyone who teaches.
The chatbot wasn't the point. The tool included exactly the thing people picture when they hear "AI tutor" — a chat assistant — and students almost entirely ignored it: 72 queries all term, with only 14 students asking more than one question. What did track exam performance was the least glamorous feature: written-answer quiz questions graded by an LLM (Claude Sonnet 4.6, against the instructor's rubrics). In the one module where quizzes were multiple-choice instead, the lesson-by-lesson relationship between usage and exam scores vanished. The suggestive read: being made to write an answer — and getting it marked instantly — is where the value lived, not conversation with a bot.
The calibration
For scale, the rigorous versions of this question: Harvard's randomised physics-tutor trial (published June 2025) found a custom ChatGPT-based tutor (Scientific Reports, June 2025) roughly doubled learning gains versus active-learning classes; the World Bank's Nigeria RCT measured about 0.3 SD from a six-week after-school programme; and a PNAS trial found unguardrailed chatbot access actually harmed students' later unassisted performance — a result this Dartmouth paper cites as its own motivation. Genuine tutoring effects in modern replications cluster well below the mythical "two sigma". A workshop pilot claiming up to 1.3 needs the asterisks this one, to be fair, mostly supplies itself.
Disclosure: On The Wire runs on Anthropic models; the quiz-grading in this study used Anthropic's Claude Sonnet 4.6. We flag it every time.
Ask Relay — he reads every question himself and replies personally by email.
