Does AI help or hurt learning? It depends how it’s used.

The Economist recently highlighted a striking study of 26,811 pupils in China. Students using general-purpose AI did their homework faster and scored better on it - but later scored 20% worse in exams.

That raises an important question: is the effect about AI, or about how students use it and what the tool asks them to do?

We looked at what happened in our own study, using data from actual 2026 GCSE exam results. Inkling users scored 7% above non-users, rising to 15% above among pupils who used it for five hours or more.

−20%

Exam score after adopting general-purpose AI

Strömberg, Lei & Wu

+7%

Exam score for pupils who used Inkling

vs non-users

+15%

Exam score for pupils who used Inkling for 5+ hours

vs non-users

What the charts show


At headline level, the contrast is stark. In the Stromberg study, pupils using general-purpose AI scored 20% below non-users. In ours, Inkling users scored 7% above non-users, rising to 15% above among pupils who used it for five hours or more.

Exam performance relative to non-users: general-purpose AI at minus 20 percent; Inkling any use plus 7 percent, three hours or more plus 12, five hours or more plus 15.

The Inkling figures above are raw group averages. Adjusting for pupils’ prior attainment gives +7%, +12% and +13% respectively.

Exam score distributions in the two studies. In the China study the AI users' curve sits left of 100; in the Inkling study the users' curve sits to the right of it.

The exam-score distributions show the same pattern. In the Stromberg study, pupils using general AI shifted towards lower scores; in ours, Inkling users shifted towards higher ones.

Both charts use the same horizontal scale. Each vertical curve is simply the distribution within that study.

Final exam score against earlier performance. In the China study the relationship breaks down at the highest homework marks; in the Inkling study it holds, with users above non-users.

The relationship with prior attainment also looks very different. In the Stromberg study, the relationship broke down: pupils with the highest AI-assisted homework marks went on to perform poorly in the exam.

In our study, January mock results continued to predict the final exam. At the same mock score, Inkling users finished around 3 percentage points higher than non-users (p = 0.013).


Why might the results be so different?

When AI does the work, learning can suffer

Strömberg, Lei and Wu found an unusual pattern. Once pupils started using generative AI, their homework marks went up and they completed it faster. But their exam performance fell.

The effect was strongest among pupils who combined very high homework marks with very short completion times - consistent with using AI to outsource the work rather than help them learn it. Pupils who kept spending similar amounts of time on their homework saw much less of a penalty.

The problem may not be AI itself, but what happens when the technology removes the thinking the homework was meant to produce.

A tutor keeps the thinking with the pupil

Inkling is designed to work differently. It checks what a pupil understands, asks questions rather than simply giving answers, and gives more help only when they need it. As they improve, the support falls away.

In our study, the pattern in the final exam went the other way: pupils who used Inkling scored higher than those who did not, and heavier users tended to do better still.

The aim is the same as with a good human tutor: give the pupil enough support to make progress, while leaving the intellectual work with them.

“AI in education” is not one thing

It is tempting to talk about whether “AI in education” works.

But that groups together products that ask very different things of the student.

An answer machine can make schoolwork easier by doing more of it for them.

A tutor should make learning more effective by helping them do more for themselves.

Our results suggest that distinction matters.

___

A note on the comparison

These are separate studies, not a head-to-head trial. The Stromberg study is much larger: 26,811 pupils across multiple subjects and 30 months. Our study covers 113 pupils, one school, one subject and eight months.

Neither study is a randomized trial. Strömberg et al. use a large 30-month difference-in-differences study; ours follows pupils over eight months, with quasi-random variation in Inkling exposure within the school.

We also do not have homework completion times, so we cannot directly test the mechanism identified in the Stromberg study.

The numbers should therefore not be read as −20% versus +15% in a controlled comparison. What they show is simpler: two very different uses of AI were associated with very different patterns of learning.

Sources:
Strömberg, D., Lei, V. & Wu, Y. (2026), “ The Generative AI Learning Penalty: Evidence from Chinese Secondary Education”, SSRN working paper; covered by:
The Economist, "Does AI stop children from learning?", 18 August 2026.
Inkling figures from the Inkling 2026 efficacy analysis and production data.