English

Using AI to Study but Grades Won’t Improve? 5 Mistakes to Fix

92% of students now use generative AI, yet grades stay flat. Peer-reviewed research points to five mistakes in using AI to study — and how to fix them.

StudyVerso Editorial 5 min read
Using AI to Study but Grades Won’t Improve? 5 Mistakes to Fix


Some 92% of UK undergraduates now use generative AI in their studies, up from 66% a year earlier, according to a Higher Education Policy Institute (HEPI) survey of 1,041 students published in February 2025. Grades have not moved with adoption. A randomized trial published in PNAS in June 2025 suggests why: students who practiced math with unrestricted GPT-4 scored 17% worse on later exams than peers who never touched it. The gap between using AI to study and actually learning with it has become one of education’s defining puzzles.

The distinction matters for the millions of students opening a chatbot next to their notes this academic year. The evidence increasingly shows that outcomes depend less on whether students use AI and more on how. Five specific usage patterns, documented across peer-reviewed studies and large-scale surveys, separate the students whose grades improve from those whose grades quietly erode.

📊 Claves rápidas

  • 92% of UK students used generative AI in 2025, up from 66% in 2024, according to HEPI.
  • Students with unrestricted GPT-4 access solved 48% more practice problems but scored 17% lower on unassisted exams, per a 2025 PNAS trial.
  • A tutor-style version of the same model raised practice scores by 127% without damaging exam performance.
  • An MIT Media Lab experiment reported that 83% of AI-assisted writers could not quote their own essay minutes after finishing it.

Record Adoption, Flat Outcomes: The Context

Generative AI reached near-universal adoption among students in under three years. According to HEPI (2025), 92% of UK undergraduates use AI tools and 88% have used them for assessed work, up from 53% in 2024. No comparable, documented rise in academic performance has accompanied that surge.

The pattern extends beyond the UK. Anthropic’s Education Report, published in April 2025 and based on roughly 574,000 anonymized student conversations with its Claude models, found that a large share of exchanges were «direct» requests: students asking for answers or finished output with minimal back-and-forth.

That usage profile collides with decades of cognitive science. Learning gains come disproportionately from effortful retrieval — struggling to recall and apply material — rather than from re-reading polished explanations. When a chatbot absorbs the struggle, it can also absorb the learning.

The Crutch Effect: Why Using AI to Study Can Lower Grades

A randomized trial with nearly 1,000 Turkish high-school students, published in PNAS (2025) by researchers at Wharton and the University of Pennsylvania, found that unrestricted GPT-4 access lifted practice-session scores by 48% — then cut unassisted exam scores by 17%. A guardrailed tutor version of the same model eliminated the harm.

The study isolates the first two mistakes on the list. Mistake one: asking for answers instead of explanations. Students in the unrestricted arm frequently requested solutions outright, performed well during practice, and then collapsed when the tool was removed.

«These results suggest that students attempt to use GPT-4 as a ‘crutch’ during practice problem sessions, and when successful, perform worse on their own.»

— Hamsa Bastani and co-authors, «Generative AI Without Guardrails Can Harm Learning,» PNAS (2025)

Mistake two follows directly: using a default chatbot when a tutor mode exists. The trial’s second arm, «GPT Tutor,» gave hints instead of solutions. It boosted practice performance by 127% and left exam scores intact. The interface, not the underlying model, made the difference.

ConditionPractice performanceUnassisted exam
No AI (control)BaselineBaseline
GPT Base (answers on demand)+48%−17%
GPT Tutor (hints, no solutions)+127%No significant harm

Skipped Self-Testing, Unchecked Output and the Illusion of Mastery

Three further mistakes compound the crutch effect. HEPI (2025) found 18% of students had pasted AI-generated text directly into assessed work, often unverified. And an MIT Media Lab preprint (2025) reported that 83% of participants writing essays with ChatGPT could not accurately quote their own text minutes later.

Mistake three is replacing self-testing with AI summaries. Research on study techniques, including Dunlosky and colleagues’ widely cited 2013 review, consistently ranks practice testing among the most effective methods and passive summarizing among the weakest. A chatbot that condenses forty pages into ten bullet points feels efficient. It also removes the retrieval effort that makes material stick.

Mistake four is trusting output without verification. Language models still fabricate citations, dates and formulas. Students who memorize an unchecked explanation risk internalizing a fluent error — and in HEPI’s data, the fear of «false or biased results» already ranks among students’ top concerns.

Mistake five is mistaking fluency for mastery. The MIT experiment, which used EEG to compare writers working with and without ChatGPT, found lower neural engagement and weaker recall in the AI-assisted group — a phenomenon its authors call «cognitive debt.» The preprint has not yet completed peer review, but its recall finding echoes the PNAS exam results. Tool choice shapes these habits too, as comparisons such as ChatGPT for Teens vs Gemini Study Notebooks illustrate: products differ sharply in how much struggle they preserve.

MistakeEvidenceDocumented fix
Asking for answers, not explanationsPNAS (2025): −17% on examsRequest hints and worked reasoning
Using default chat instead of tutor modesGPT Tutor arm: no exam harmEnable study/tutor modes
Summaries instead of self-testingDunlosky et al. (2013)Generate quizzes, not digests
Not verifying outputHEPI (2025): 18% paste AI textCross-check against course sources
Confusing fluency with masteryMIT Media Lab (2025): 83% recall failureReproduce material unassisted

What the Evidence Means for Students and the EdTech Sector

The research points to a design problem as much as a study-habits problem. In 2025, OpenAI launched Study Mode, Google added Guided Learning to Gemini and Anthropic shipped a Learning Mode for Claude for Education — all restricting direct answers in favor of Socratic questioning, mirroring the PNAS tutor condition.

For students, the practical reading is narrow but actionable: the same model that depresses grades in one configuration improves them in another. The variable is whether the tool preserves effortful practice or replaces it.

For the sector, the incentives are less aligned. Answer-on-demand products retain users better than tools that make them work. EdTech apps — from Khanmigo to Spanish startups such as Modo Cheto — now face the same tension: measurable learning versus frictionless convenience. Universities, meanwhile, are redesigning assessment around the assumption that 88% of students already bring AI to their coursework.

Arturo P.L. — Arturo P.L. cubre inteligencia artificial aplicada a la educación en StudyVerso. Ingeniero, ex-consultor y co-fundador de una startup EdTech. Analiza lanzamientos de modelos, políticas universitarias y adopción real de IA en aulas españolas y LatAm.

The open question is no longer whether students will keep using AI to study — adoption figures have settled that. It is whether tutor-style guardrails become the default before another cohort discovers, at exam time, the difference between borrowed fluency and earned understanding. The 2026 round of surveys, already underway at HEPI, will offer the first read.

Estadísticas verificadas hoy contra fuentes primarias antes de redactar:

– HEPI, [Student Generative AI Survey 2025](https://www.hepi.ac.uk/reports/student-generative-ai-survey-2025/) — 92% de uso (66% en 2024), 88% en trabajos evaluados, 18% pega texto de IA, n=1.041.
– Bastani et al., [*Generative AI without guardrails can harm learning*, PNAS 2025](https://www.pnas.org/doi/10.1073/pnas.2422633122) — +48%/−17% (GPT Base), +127% sin daño (GPT Tutor); la cita del blockquote es literal del abstract.

El dato del MIT Media Lab (83%, «cognitive debt») es un preprint de 2025 y así lo señalo en el cuerpo. La FAQPage la omití: la pieza no tiene preguntas estructuradas reales.

Avatar de StudyVerso Editorial
StudyVerso Editorial