English

Using AI to Study Won’t Boost Your Grades: What You’re Doing Wrong

Using AI to study rarely improves grades. Research from Wharton, MIT and Anthropic shows why passive prompting backfires — and what actually works.

StudyVerso Editorial 5 min read
Using AI to Study Won’t Boost Your Grades: What You’re Doing Wrong


Ninety-two per cent of UK undergraduates now use generative AI, according to a Higher Education Policy Institute (HEPI) survey of 1,041 students published in February 2025. Their grades, however, are not following. A field experiment led by Wharton and University of Pennsylvania researchers, published in 2024, found that students using AI to study with a standard chatbot scored 17% worse on exams than peers who practiced without it. The gap between adoption and results is becoming the central tension of the AI-in-education story.

The finding matters because it cuts against the default assumption of millions of students: that more AI equals better outcomes. The emerging evidence says the tool is not the variable — the usage pattern is. Students who treat chatbots as answer machines learn less. Students who treat them as tutors hold their ground or improve.

📊 Claves rápidas

  • According to HEPI (2025), 92% of UK undergraduates use generative AI, up from 66% in 2024.
  • A Wharton-led experiment (2024) found unrestricted ChatGPT access lowered exam scores by 17%.
  • The same study showed a guardrailed «tutor mode» eliminated the exam penalty entirely.
  • Anthropic’s Education Report (2025) found roughly half of student AI conversations were direct answer-seeking.

The Studies Behind the AI Grades Paradox

Access to a standard ChatGPT interface lowered final exam scores by 17% among nearly 1,000 high-school students, according to a 2024 field experiment by Wharton and University of Pennsylvania researchers. Practice-session performance rose sharply while the chatbot was available. The gains reversed the moment it was taken away.

The experiment, run across math classes in Turkey, split students into three groups. One practiced with a standard GPT-4 chatbot, one with a «GPT Tutor» configured to give hints rather than answers, and one with no AI at all. The standard chatbot group solved 48% more practice problems correctly. On the closed-book exam, they underperformed the no-AI group.

«These results suggest that students attempt to use GPT-4 as a ‘crutch’ during practice problem sessions and, when successful, perform worse on their own.»

— Bastani et al., «Generative AI Can Harm Learning», University of Pennsylvania / Wharton working paper, 2024

Neuroscience is starting to sketch the mechanism. An MIT Media Lab study published in June 2025 monitored 54 participants with EEG while they wrote essays. Those using ChatGPT showed the weakest brain connectivity of any group, and a large majority could not accurately quote from essays they had submitted minutes earlier. The researchers labeled the pattern «cognitive debt»: output produced, learning deferred.

Why Using AI to Study Often Backfires

Cognitive science offers a consistent explanation for why using AI to study fails to lift grades: retrieval practice. Decades of research, including landmark 2006 experiments by Roediger and Karpicke, show that actively recalling information strengthens memory far more than re-reading it. Copying a chatbot’s answer involves no retrieval at all.

Exams reward the ability to reconstruct knowledge under pressure, without assistance. A chatbot that hands over polished solutions removes exactly the struggle that builds that ability. The student experiences fluency in the moment and mistakes it for mastery. Researchers call this the illusion of competence, and AI industrializes it.

There is a second, quieter problem: error propagation. In the Wharton experiment, students frequently asked the chatbot for the final answer rather than the method. When the model got arithmetic wrong, students copied the mistake without noticing. Delegating the reasoning also means delegating the quality control. Many of these behaviors are correctable, as a companion analysis of five fixable mistakes students make when studying with AI details.

Answer Engine or Tutor: Where AI Study Habits Diverge

Roughly half of student conversations with Claude were «Direct» — requests for answers or finished content with minimal back-and-forth — according to Anthropic’s Education Report, which analyzed 574,740 anonymized university conversations in 2025. The other half resembled dialogue: students iterating, asking follow-ups and requesting explanations, a pattern much closer to tutoring.

That split maps neatly onto the experimental evidence. The Wharton study’s «GPT Tutor» condition — a version prompted to give hints, withhold final answers and flag mistakes — boosted practice performance by 127% while producing no measurable exam penalty. Same model, same students, opposite outcome. The interface design, not the underlying AI, decided whether learning happened.

The broader literature points the same way. A 2025 meta-analysis of 51 studies published in Humanities and Social Sciences Communications found ChatGPT can significantly improve learning performance, but chiefly when embedded in structured, instructor-designed activities. Unstructured, self-directed use is where the negative results cluster. Using AI to study is not one behavior; it is a family of behaviors with divergent outcomes.

Usage patternTypical behaviorDocumented outcome
Answer-seekingPasting problems, copying solutions17% lower exam scores (Wharton, 2024)
Essay delegationAI drafts the written workWeakest neural engagement; poor recall of own text (MIT, 2025)
Guardrailed tutoringHints and step-by-step feedback, no final answers127% practice gain, no exam penalty (Wharton, 2024)
Self-testing with AIGenerating quizzes, checking recalled answersConsistent with retrieval-practice gains (Roediger & Karpicke, 2006)

What the Evidence Means for Students and Universities

For universities, the findings land at an awkward moment. According to HEPI (2025), 88% of UK undergraduates already use generative AI in assessments, up from 53% a year earlier. Institutions can no longer choose between adoption and abstention. The open variable is whether students use these tools as tutors or as answer machines.

For students, the practical translation is uncomfortable but clear. The productive uses of AI are the effortful ones: asking for a hint instead of a solution, requesting a quiz instead of a summary, explaining an answer back and asking the model to find the holes. The EdTech market is slowly converging on this, with tools from Khanmigo to European startups such as Modo Cheto building products around guided practice rather than answer delivery.

For institutions, the burden shifts from policing to design. Detection tools remain unreliable, and the HEPI data suggests prohibition has already lost. The alternative being tested at several universities is assessment redesign: more invigilated components, more oral defenses, and explicit instruction in AI study technique — treating it as a skill to be taught rather than a temptation to be banned.

The Question Heading Into Exam Season

The research base is young, the models change quarterly, and most of the strongest evidence comes from a handful of settings. What is already hard to dispute is the direction: using AI to study pays off only when the student stays in the loop of effort. Whether universities can redesign assessment faster than students default to the path of least resistance is the question the 2026-27 academic year will start to answer.

Isabel A.M. — Isabel A.M. escribe sobre pedagogía, métodos de estudio y el impacto de la tecnología en la vida del estudiante. Co-fundadora de una startup EdTech, sigue de cerca el sector universitario, las oposiciones y las certificaciones de idiomas.

Avatar de StudyVerso Editorial
StudyVerso Editorial