part of EQUORA Institute
HUEN
EQUORA Institute · Active research thread · NeverNormal
Answers are Questioned™ · Cognition & AI · Active research strand

Should schools restrict AI? Where the limits go is what matters.

Reading the 2025 PISA report, the conclusion seems self-evident: students who use AI for learning score lower, so schools should allow less AI. The experimental data, however, show something more specific: depending on when and how AI tools are used, they can improve results and they can worsen them.

Reading · 11 min
Status · Open question
Frame · Sequence and prerequisite
01 · The number

A school year, in a single chart

PISA 2025 was published on 9 September, and one number travelled out of its AI chapter. Fifteen-year-olds who use an AI chatbot almost daily to draft written assignments averaged 481 points in science; those who never do, 509. The 28-point gap survived adjustment for socio-economic background, and by the report's own conversion it equals roughly a year of schooling.

The head of PISA compressed the reading into one line at the launch: nobody gets fit by watching sports. The international teachers' federation said AI is being pushed as a magic solution while children outsource their thinking to a bot. The infographics put the one-school-year gap between daily users and non-users on the front page. Out of these three elements, the answer this page questions assembles itself: AI harms learning, so schools should restrict it.

02 · What PISA sees

A snapshot, self-reported

PISA is cross-sectional: it asks students at one point in time how much AI they use, and tests what they know at the same moment. The OECD says so itself in the report: students who use AI may differ from non-users in prior achievement, motivation, support and study habits, so the association describes a state, and a causal direction is beyond what this data can carry.

The same chapter holds two less-quoted details. One is the shape of the curve: for use “to help me learn”, weekly users perform at the same level as non-users, while daily users and once-or-twice-a-year users sit below both. That is an inverted U, and reading a “less is better” rule out of an inverted U is exactly as well founded as reading “more is better” out of it. The other detail is Box I.4.2: where schools give students classroom practice in judging the quality of AI-generated content, frequent AI users perform slightly better than non-users. The same box records that disadvantaged students receive this practice less often.

Inverted U — a relationship between amount of use and outcome where the middle is best and both extremes are weaker. PISA found the same shape for digital-device use at school; for AI, weekly frequency sits at the top. On its own the curve says only that something other than quantity decides.
03 · The experiment

Same tool, two outcomes

PISA sees what happens. A randomised experiment also sees what would happen if it were done differently. The cleanest such study appeared in PNAS in 2025: nearly a thousand Turkish high-school students in maths classes were randomly split into three groups. The first practised without AI. The second got a standard ChatGPT-like interface on GPT-4. The third got the same model, but its prompt contained the solution to each problem and an instruction: help step by step, hold back the full solution. After practice, everyone sat the same exam, without any tool.

Bastani et al., 2025 · ~1,000 high-school maths students · difference from the control group
No AI (control)
0
Unrestricted GPT-4 (“GPT Base”)
+48%
Guided tutor (“GPT Tutor”)
+127%
Same model, same students, same material. The guided variant had the solution in its prompt but gave the student only the next step and hints. On the exam, the difference between the guided group and the control group is statistically zero.

The toggle is the axis of this whole page. During practice both AI groups are dramatically better than the control, and the guided tutor is the better of the two. On the exam, the unrestricted group is 17% worse than the group that never saw AI — and the guided group sits level with the control. In the authors' words: without guardrails, students used the tool as a crutch and then performed worse on their own; the guided variant largely erased that effect.

Two further experiments give the other half of the picture. The World Bank's six-week programme in Nigeria, in which Copilot was used under teacher supervision on assigned material, produced a gain of 0.31 standard deviations on the overall test and 0.23 in the main subject, English; the authors convert its cost-effectiveness into one and a half to two years of business-as-usual schooling. In Harvard's introductory physics course, a purpose-built AI tutor beat the active-learning classroom by more than double the median learning gain, in less time. The same model family appears in all three experiments. The outcome turned on interface design and on the moment of use.

Guardrail — the set of instructions given to the model that configures it to hold back the answer and steer towards the next step. In the Bastani experiment it had three parts: the solution in the prompt (against hallucination), a ban on giving away the full solution, and a list of common student errors. The limit sits on the model's behaviour, well away from access.
04 · The sequence

Estonia taught first, then let it in

There is an odd pair on the PISA country list. Estonia sits in Europe's top group in science, and its students are among those using AI for schoolwork least often. The same country launched the AI Leap (TI-Hüpe) programme in August 2025: all 154 upper-secondary schools take part, around 20,000 students in grades 10 and 11, close to 5,000 teachers. Students get a ChatGPT-based application designed to guide thinking rather than hand over answers, replying only in Estonian; teachers get premium access and professional learning communities; the University of Tartu and Stanford track learning outcomes.

Side by side, the two facts describe a sequence. PISA data collection closed in spring 2025 and the programme started in August: Estonia built reading and reasoning performance first, and let the tool in afterwards — with a plan, teacher preparation and a research programme. The first-year process figures are in (94% of participating teachers used AI in their work, 63% redesigned their teaching); learning outcomes begin arriving from late 2026. This programme is the only place in the world where the question “how long does the skill take to build” is generating national-scale data.

05 · The prerequisites

Three skills, before and alongside the tool

If the outcome turns on the moment of switching on, the question becomes what should be in place by then. Three skills stand out, each with a research anchor. Their foundation is built before AI; their development continues alongside guided AI use, as the Nigerian programme and the guided tutor show — it is answer-generating use that hides their absence.

1 · Critical evaluation
Practised habit of judging the quality of an AI answer. PISA Box I.4.2: where this is exercised in class, the frequent user gains an edge. The Microsoft–CMU knowledge-worker survey (Lee et al., 2025) shows the same from the other side: the higher the confidence in the tool, the less critical thinking is spent.PISA 2025, Box I.4.2 · Lee et al., CHI 2025
2 · Delegation and accountability
Specifying a task precisely and checking the result — the same skill whether the executor is a person or a machine. The central observation of the Bastani experiment: students accepted correct and incorrect AI answers alike. Delegation and verification form one skill; apart, neither works.Bastani et al., PNAS 2025
3 · Abstraction and modelling
Thinking in models and scenarios — so there is something to hold the machine's output against. Kalyuga's expertise reversal effect: the same support that helps a novice hurts an advanced learner, so the rule for AI use differs by level, and the level of abstraction signals the switch.Kalyuga et al., 2003 · Sweller, 1988

What the three share is that each was valuable before AI and stays valuable after it. The difference is that their absence used to surface slowly and now surfaces at once: an answer-generating tool hides the missing skill during practice and exposes it on the exam. The toggle in 03 shows exactly that.

06 · The dip

The curve goes down first

This page has a sibling: Are we getting dumber because the technology is getting smarter? It sets out a three-mechanism model whose first step belongs here. At the start, AI overloads: it hands the learner a new task (judging, comparing, checking) for which no schema exists yet. While the schema builds, performance drops, and only afterwards turns upward. That also explains one of PISA's puzzles: the most intensive, earliest adopters are deepest in the dip.

Nobody has data on the length of the dip today. A school rollout that follows the principle of gradualism should budget a three-, six- or twelve-month learning curve, and that figure is for now a hypothesis — the Estonian programme will be the first place where it becomes testable. PISA's single snapshot sees one point on that curve, on the downward stretch. Turning that one point into a “restrict it” rule stops the learning at the bottom of the learning curve.

07 · The twist

Where the limit should go

Put the four pieces together — the inverted U, the toggle, the Estonian sequence and the dip — and “restrict AI” takes a more precise shape. A limit is needed — in three places, and all three sit elsewhere than where the debate currently draws it.

First, on the moment. Bjork's “desirable difficulties” and Kapur's “productive failure” research show that an unaided attempt before help yields durable knowledge. The Bastani tutor's “hint, no answer” rule is a technical form of this; the classroom version is one sentence: AI switches on after your own first draft. Second, on the measurement. Performance during practice misleads — both AI groups were at their best there. The only valid success criterion is unaided performance after the AI period. Anyone designing a school programme should carry that single testing principle in and drop the “how many students use it” metrics. Third, on the prerequisite. The foundation of the three skills from 05 enters the curriculum before the tool; their further development runs alongside guided use, and answer-generating use waits until the checking skill is working.

And here the equity argument turns over as well. PISA's least-quoted finding is that disadvantaged students receive AI-evaluation practice less often. If the prerequisite skill decides the outcome, then a school ban protects the students whose skill comes from home — while the rest use the tool behind the ban anyway, only without preparation. In this reading, teaching AI literacy is equality policy, and prohibition locks in the existing gap.

The reframe
The limit sits well on the tool's behaviour and on the moment it is switched on. Placed on access, it halts learning at the bottom of the learning curve.
The next question

If the outcome is decided by the moment of switching on and by the prerequisite skill — how long does the skill take to build, and who pays for the dip until then?

Nobody has run this study to the end: the Bastani experiment sees one subject in one school year, the Nigerian programme six weeks, PISA one spring. The Estonian programme is the first place where, over years and on a national sample, it will show whether the curve turns upward and how long that takes. The prediction that follows from this page: the final outcome is decided by when the tool is switched on relative to the attempt, by what is in place by then, and by what is measured afterwards. That is the study waiting to be done.

Let's build the research →
Sources. OECD (2026). PISA 2025 Results (Volume I), Executive Summary and Chapter I.4 “Students' use of artificial intelligence for schoolwork”, Box I.4.2; doi:10.1787/73451bc5-en · Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö. & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS 122(26); doi:10.1073/pnas.2422633122 · De Simone, M. et al. (2025). From Chalkboards to Chatbots. World Bank Policy Research Working Paper 11125 · Kestin, G. et al. (2025), AI tutor vs. active learning, randomised trial, Harvard physics · Lee, H.-P. et al. (2025), knowledge-worker survey on generative AI and critical thinking, CHI 2025 · Kalyuga, S. et al. (2003), the expertise reversal effect · Sweller, J. (1988), cognitive load theory · Bjork, R. A., “desirable difficulties” · Kapur, M., “productive failure” · AI Leap / TI-Hüpe Foundation (2026), first-year report, tihupe.ee · Schleicher, A., PISA launch and Euronews interview, 10 September 2026 · Education International, press release, 9 September 2026. Figures paraphrased; see the originals for exact values.
Research provenance
This page comes out of research at the EQUORA Institute and captures one state of that work rather than a settled institutional position. That state rests on the findings available at the time of publication; later findings appear here only where the page has been updated, which the date shows. AI takes part throughout the research process as a thinking partner; responsibility for interpretation and publication remains human.
Published: 15 September 2026
Papp László · EQUORA InstituteHow We Research →