The Loop  ·  Issue 031

The Loop

A field journal of the AI frontier — for engineers who ship.

§ News

By AI Blog Editor
Jul 6, 2026 · 17 min read

Homework up eighteen, exams down twenty-four — the 26,811-student panel that measured the AI learning penalty on a two-year fuse

A June 2026 CEPR paper tracking 26,811 Chinese secondary students over 30 months finds AI adoption raises homework scores 18% but drops monthly exam scores 20% within six months — and cuts high-stakes entrance-exam scores 18–24% on a roughly two-year fuse.

A colour photograph of Chinese secondary-school students walking out through the front gate of Beijing 101 Middle School on June 7, 2021, after finishing the first paper of the National College Entrance Examination — the gaokao. Groups of parents in the foreground wait to greet them. A June 2026 CEPR discussion paper by David Strömberg, Victor Lei and Yanhui Wu tracks 26,811 Chinese secondary students across 30 months and finds that generative-AI adoption raised homework scores by 18% but lowered closed-book exam scores by 20% within six months, with a two-year fuse of 18–24% on entrance-exam performance.
Examinees leaving Beijing 101 Middle School after the first paper of the 2021 gaokao. Photograph by N509FZ, CC BY-SA 4.0 via Wikimedia Commons.

On June 19, 2026, a CEPR discussion paper by David Strömberg, Victor Lei and Yanhui WuDP21577, also on SSRN as #6868618 — put the sharpest number to date on a claim that the ed-tech industry has been dancing around for two years. Across 26,811 Chinese secondary students, grades 7 through 12, tracked over 30 months of panel data in a county of more than a million people, generative AI raised homework scores by 18% and cut completion time from 64 minutes to 45. Over the same period, monthly closed-book exam scores fell 20% within six months, and high-stakes entrance-exam scores dropped 18 to 24%, with the full penalty surfacing only after roughly two years.

That is the AI learning trap, measured, timed, and priced. It is the first difference-in-differences study on the tape big enough, long enough, and specific enough that the ed-tech vendors selling AI homework helpers cannot wave it away.

Homework up, exams down

The paper's headline finding is the split between the two things AI touches. Homework is a graded output; the metric is the number written at the top of the page. Exams are a graded input; the metric is what the student can produce with no tools in the room. AI closes the first metric and opens the second, at the same time, in the same students, on the same subjects.

Per The Decoder's write-up of the paper on July 4, homework performance went in one direction and closed-book performance went in the other. Homework scores climbed 18%. Completion time dropped roughly 30% — from 64 minutes on the average assignment to 45. Monthly closed-book exam scores dropped 20% in the first six months of AI use.

Per the Psychology Today piece by Soren Kaplan on June 25, the entrance-exam penalty — the number the students, their parents, and the Chinese university system care about — arrived on a two-year lag. Students who adopted AI in ninth grade did fine on their homework for two full years, then walked into their high-school entrance exam and scored 18 to 24 percent lower than the matched control cohort who did not.

Ethan Mollick of Wharton, who has spent two years telling anyone who will listen that the mode of AI use is what determines the learning outcome, gave the Digg summary the sentence the ed-tech marketing decks will not print. "Off-the-shelf chatbots give you the answer and undermine learning," Mollick said. "AI tutoring in support of classes is good, using AI to help with homework is bad."

That is a distinction the paper's methodology can defend and a distinction almost no consumer AI product enforces on itself.

The mechanism has a name and a percentage

The finding that carries the paper is not that AI use hurts learning on average. It is that the harm is concentrated in a specific behaviour, and the behaviour has a signature the panel data caught cleanly. Roughly 80% of AI-using students show what the authors call homework outsourcing — assignment completion times that are far shorter than the pre-AI baseline, paired with homework scores that are much higher. The other 20% show completion times close to normal, use AI more like a reference or a tutor, and lose almost none of the closed-book gains.

That 80/20 split is the whole ballgame. It says the population is bimodal: most students use AI to replace the effort, and a minority use it to supplement the effort, and the exam data cleanly separates the two. It is also the reason the effect is so large in aggregate. If 80% of your treatment group is outsourcing the practice reps, of course the closed-book scores fall — the students never did the practice.

The subject breakdown maps onto this in a way that is not surprising once you see it. Social sciences lost 27%. STEM lost 22%. English lost 17%. Chinese lost 9%. The subjects where the assignment is write a coherent argument or explain a mechanism in your own words — the ones where an LLM has the largest edge over a student — are the subjects where the students who outsource lose the most, because those are also the subjects where the practice of writing the argument is what the exam actually tests.

Chinese-language work loses the least because a native speaker cannot outsource fluency to a chatbot in a language they already dream in. Social science loses the most because a chatbot can write a passable essay about the Warring States period in about eleven seconds, and a fourteen-year-old who has spent two years accepting those essays as their own work cannot write one at all.

The two-year fuse

The temporal shape of the penalty is the finding that ed-tech buyers should stare at longest. Within six months of AI adoption, closed-book exam scores are already down 20%. Within two years, the entrance-exam penalty is at its full 18–24%. In between, the students look, on paper, like they are doing better — better grades, faster turnaround, happier teachers, calmer parents.

That is a two-year window during which the observable signal points up and the underlying capability points down. It is exactly the window in which an ed-tech vendor can run a paid pilot, publish the homework-score number, and exit the district before the entrance-exam number comes in. It is also the window in which an anxious parent buys the AI tutor subscription, sees the grades climb, and never learns that the ninth-grade year the tutor covered was the year the child stopped learning to write.

A colour photograph of Chinese secondary-school students seated at desks in an examination hall, reviewing notes on the morning of the 2018 gaokao at Beijing No. 4 Middle School. The image shows the intensity of preparation typical of the Chinese secondary-school system. The Strömberg, Lei and Wu panel, tracking 26,811 students in grades 7 through 12 over 30 months, isolates a two-year lag between generative-AI adoption and the full impact on high-stakes entrance-exam performance — a window during which homework grades climbed while underlying capability eroded.

Where the losses concentrate

Two demographic findings in the paper are worth flagging because they invert the intuition that ed-tech pitches lean on. First, the losses are largest among high-achieving students, not struggling ones. High achievers had the most to gain from doing the homework properly and the most cognitive slack to outsource. They took the outsourcing option in bulk, and the exam data caught them. Second, boys lost more than girls. The paper does not fully explain the sex gap, but the pattern is consistent with the outsourcing-signature framing — boys in this population appear to hit the finish faster and move on setting on the AI tool more than girls do.

The paper also finds that junior students — grades 7 and 8, the youngest in the panel — lose more than older students. That is the group most schools are least worried about, because the entrance exam is years away and the immediate grades look fine. It is also the group whose foundational years of writing, arithmetic, and reading practice are the ones the outsourcing behaviour is eating.

The buried finding, in other words, is not AI hurts weak students. It is AI hurts the students you thought were doing best, the youngest, and the boys, and it hurts them in the subjects where the essay is the whole point.

The DeepSeek anchor

The paper's adoption curve gives a rare, dated anchor for when generative AI became a normal thing for a Chinese teenager to use for homework. Adoption in this county rose from near zero to about 80% over the study period. The major spike coincides with two dates: DeepSeek V2.5 in September 2024 and DeepSeek R1 in January 2025. The Chinese-language, free-at-point-of-use, phone-friendly, reasoning-capable model landed and the adoption curve went vertical.

That is the piece of the story a US ed-tech operator should read twice. The Chinese cohort in this study is not using a US chatbot behind a VPN. It is using a domestic model that runs on domestic app stores, in the domestic language, with domestic curricula in its training data. The homework-outsourcing behaviour scales the moment the tool stops feeling foreign. The equivalent curve in the US and Europe is dragging behind by roughly the amount of friction that language, payment, and app-store availability still add — and that friction is dropping every quarter.

What this means

  1. The AI homework trap is measurable, and the fuse is two years long. The Strömberg-Lei-Wu numbers turn the classroom-teacher intuition into a difference-in-differences result on 26,811 students. Any school district piloting an AI homework tool that publishes a six-month grades-up number without a matched closed-book exam number is, at best, publishing half the study — and, at worst, publishing the half that inverts within twenty-four months.

  2. The 80/20 split is the policy handle. The paper does not say AI hurts learning. It says roughly four in five student AI users hurt their own learning by outsourcing, and one in five use AI in a way that does not. That is a distinction schools can teach, product designers can enforce, and vendors can build against. It is also the distinction almost no consumer AI product currently designs for, because finish faster and move on is the growth metric.

  3. High-achievers, boys, and juniors are the buried headline. Ed-tech pitches have spent two years telling parents of struggling students that AI will help their kid catch up. The population that actually paid the largest penalty in this data was the high-achievers, and the demographic slice that outsourced hardest was young boys in grades 7 and 8. That is the exact profile the AI study buddy subscriptions target on TikTok.

  4. DeepSeek's launch curve is the timeline every Western district will get in the next twenty-four months. The Chinese cohort's adoption went vertical after DeepSeek R1 dropped in January 2025. The Western equivalent is dropping every quarter as friction disappears. Any US or European district that thinks the numbers in this paper are a China story is going to publish the same numbers on its own students by 2028.

The paper is the first entry on a shelf that is about to fill up. It is also the first entry that was rigorous enough to catch the two-year fuse. Every follow-up study that runs shorter than two years is going to conclude AI helps. The one that runs twenty-four months is going to conclude something else.

* * *

Thanks for reading. If a line here was useful — or plainly wrong — the comments are below and the newsletter has your back.

Elsewhere in this issue

3 more
  1. 01

    News

    The means of production — Palantir posted $1.94 billion in one quarter, then used the shareholder letter to accuse OpenAI and Anthropic of Marxism

    Aug 4, 2026

  2. 02

    The Patch

    The Patch — August 4, 2026

    Aug 4, 2026

  3. 03

    News

    The rug pulled, twice — Timothy Gowers wrote the mathematical-culture case against Astra six days before OpenAI announced Astra

    Aug 3, 2026

Letters

Arguments, corrections, questions. Anonymous comments allowed; be kind, be specific.