PISA Warns AI Can Hurt Your Learning: How to Use It Right
PISA researchers at the OECD warn passive AI use can hurt student learning. New data shows when AI and learning clash — and how to use chatbots right.

The OECD, the organisation behind the PISA rankings, is warning that unguided use of AI chatbots can depress student learning rather than boost it. According to a University of Pennsylvania field study (2024), high-school students with open access to GPT-4 scored 17% worse on exams than peers who studied without it. The warning lands as the OECD prepares results from PISA 2025, the first cycle to measure «Learning in the Digital World.» First findings are expected from late 2026.
The stakes are practical, not theoretical. Millions of students now open a chatbot before they open a textbook. Whether that habit builds knowledge or quietly erodes it depends on how the tool is used — and the evidence on AI and learning now points in two very different directions at once.
- Students with unrestricted GPT-4 access solved 48% more practice problems but scored 17% worse on the final exam, according to a 2024 University of Pennsylvania study.
- PISA 2022 found that 65% of students across OECD countries reported being distracted by digital devices in at least some mathematics lessons.
- A tutor-style chatbot with pedagogical guardrails eliminated the exam-score penalty observed with the unrestricted version.
- PISA 2025 will assess «Learning in the Digital World» for the first time, with results expected from late 2026.
Why PISA Is Turning Its Attention to AI and Learning
PISA, the OECD’s triennial assessment of 15-year-olds in more than 80 education systems, has tracked technology’s effect on outcomes for two decades. Its PISA 2022 report (2023) found that 65% of students across OECD countries reported being distracted by digital devices in at least some maths lessons — a finding that now frames its scrutiny of AI.
The OECD’s scepticism about screens is not new. Its 2015 report «Students, Computers and Learning» found no appreciable improvement in reading, maths or science in countries that had invested heavily in classroom technology. Andreas Schleicher, the OECD’s education director and the architect of PISA, has repeated the same caution as generative AI entered classrooms.
«Technology can amplify great teaching, but great technology cannot replace poor teaching.»
PISA 2022 added a sharper data point. Students who spent more than an hour a day on digital devices for leisure at school scored 49 points lower in mathematics than peers who did not — roughly two and a half years of schooling. The lesson the OECD draws is about mode of use, not the device itself: moderate, learning-focused use correlated with better scores.
The Evidence: When AI Helps and When It Hurts
The clearest causal evidence on AI and learning comes from a randomised field experiment with roughly 1,000 Turkish high-school students (Bastani et al., University of Pennsylvania, 2024). Students using standard GPT-4 solved 48% more practice problems correctly, then scored 17% worse than the control group on a closed-book exam.
The mechanism the researchers identified is familiar to any teacher: the chatbot became a crutch. Students copied answers instead of working through problems. Performance looked excellent while the AI was present and collapsed when it was removed. The study’s title states the finding bluntly: «Generative AI Can Harm Learning.»
The same experiment contains the counter-finding. A second group used a modified «GPT Tutor» that gave hints, asked guiding questions and refused to hand over solutions. Those students got the practice benefit — and showed no exam penalty at all. The damage was not caused by AI. It was caused by AI configured to complete work rather than teach it.
Cognitive science explains why. Learning depends on what researchers call desirable difficulty: retrieval, struggle and self-correction. An assistant that removes the struggle removes the learning. The distinction also applies to research tasks. A student who asks a chatbot to write an essay learns little; one who uses a tool like Consensus to search 220 million peer-reviewed papers and then reads and argues from the sources is doing the cognitive work themselves.
Study Modes: How AI Companies Are Redesigning Chatbots for Learning
The major AI labs have responded to the tutoring evidence with dedicated learning products. OpenAI launched Study Mode for ChatGPT in July 2025, Anthropic introduced a Learning Mode in Claude for Education in April 2025, and Google shipped Guided Learning in Gemini in August 2025 — all built to withhold direct answers.
The design pattern is consistent across the three: Socratic questioning, stepwise hints and checks for understanding replace the instant solution. It is, in effect, the «GPT Tutor» condition from the Pennsylvania experiment turned into a consumer feature. A wave of EdTech startups — from Khanmigo to Spanish apps such as Modo Cheto — is building on the same premise.
| Usage pattern | Typical behaviour | Documented effect |
|---|---|---|
| Answer engine | Student pastes the problem, copies the solution | 17% lower exam scores vs. no AI (Bastani et al., 2024) |
| Guardrailed tutor | Hints and questions only; no full solutions | No exam penalty; more practice completed (Bastani et al., 2024) |
| Research assistant | AI finds and summarises sources; student reads and writes | Preserves retrieval and synthesis effort; no causal harm documented |
The open question is whether students choose the harder mode. Study Mode and its rivals are opt-in toggles sitting next to a default chatbot that will happily write the essay. The OECD’s device data suggests self-restraint is not the norm: distraction, not enrichment, dominated when screens entered classrooms unmanaged.
What the PISA Warning Means for Students and Schools
For education systems, the practical takeaway from the OECD’s position and the 2024 experimental evidence is that AI policy should regulate the mode of use, not access itself. Blanket bans forgo a tutoring benefit; unrestricted access carries a measured 17% performance cost when the tool substitutes for thinking.
For students, the evidence supports a simple division of labour. Delegate the mechanical layer: searching literature with academic AI search engines, formatting, generating practice questions. Keep the cognitive layer: first attempts, retrieval from memory, writing the argument. Asking a chatbot to quiz you exploits the testing effect; asking it for the answer bypasses it.
For schools, PISA 2025 will bring the first internationally comparable data on how students actually self-regulate in digital environments. Universities and exam boards, meanwhile, are already shifting weight back toward supervised, AI-free assessment — the one setting where borrowed competence becomes visible.
For the EdTech sector, the tutoring evidence sets a new bar. Products that optimise for task completion now face data showing that completion and learning can move in opposite directions. The pedagogically constrained design is no longer a philosophical preference; it is the empirically safer product.
When the PISA 2025 results land, they will show for the first time how a generation raised alongside chatbots performs when the chatbot is switched off. The relationship between AI and learning will stop being a matter of position papers and become a scoreboard. The question for every student between now and then is the one the Pennsylvania data already answers: is the AI doing the work, or teaching you to do it?