The Potemkin University Trap: Using AI Without Learning Less
AI use among students hit record highs while evidence of "cognitive debt" mounts. Inside the Potemkin university trap — and how to use AI without learning less.

Generative AI has become near-universal on campus: according to the Higher Education Policy Institute (2025), 88% of UK undergraduates used AI tools for assessments, up from 53% a year earlier. At the same time, an MIT Media Lab study published in June 2025 found that essay writers relying on ChatGPT showed the weakest neural engagement of any group tested. Together, the two findings frame what researchers are starting to call the Potemkin university trap: institutions where credentials keep flowing while the learning behind them quietly hollows out.
The tension matters to anyone paying tuition or hiring graduates. If students can produce competent work without building competence, the degree still signals knowledge the holder may not have. The question facing universities this academic year is no longer whether students use AI — that debate is over — but whether they can use it without learning less.
- HEPI’s 2025 survey found 88% of UK undergraduates used generative AI for assessed work, a 35-point jump in one year.
- MIT Media Lab researchers reported in 2025 that ChatGPT-assisted writers showed reduced brain connectivity and struggled to recall their own essays.
- Anthropic’s 2025 Education Report found students most often asked Claude to create or improve academic content directly, rather than to explain concepts.
- OpenAI and Anthropic both shipped dedicated learning modes in 2025, designed to withhold direct answers and force step-by-step reasoning.
Where the «Potemkin» Framing Comes From
The term borrows from a 2025 paper by researchers at MIT, Harvard and the University of Chicago, which coined «Potemkin understanding» to describe language models that pass benchmark tests while lacking the coherent grasp of concepts the scores imply. Applied to higher education, the label describes students — and institutions — whose outputs look like learning without the substance underneath.
The original Potemkin villages were façades allegedly built to impress Catherine the Great: convincing fronts, nothing behind them. The academic analogy is uncomfortable but precise. A polished essay, a correct problem set and a passing grade have always been proxies for understanding. Generative AI breaks the proxy. It produces the artifact without requiring the cognition the artifact was designed to certify.
Universities have faced proxy failures before — essay mills, contract cheating, shared answer banks. The difference now is scale and legitimacy. AI assistance is free or cheap, instant, and in many courses explicitly permitted. The façade is no longer built in the shadows; it is built with institutional tools, sometimes under institutional licenses.
What the Evidence Says About Cognitive Debt
According to MIT Media Lab’s «Your Brain on ChatGPT» study (2025), participants who wrote essays with LLM assistance showed the lowest brain connectivity of three groups measured by EEG, and 83% could not quote from essays they had submitted minutes earlier. The authors called the effect «cognitive debt»: effort deferred now, capability lost later.
The study was small — 54 participants — and released before peer review, a caveat its lead author acknowledged openly. But its central finding aligns with decades of cognitive science. Learning is a function of effortful retrieval and struggle, what researchers call desirable difficulty. Remove the struggle and the artifact survives while the memory trace does not.
The lead researcher explained why she chose not to wait for the peer-review cycle before publishing:
«I am afraid in six to eight months, there will be some policymaker who decides, ‘let’s do GPT kindergarten.’ I think that would be absolutely bad and detrimental.»
Usage data points in the same direction. Anthropic’s Education Report, published in April 2025, analyzed anonymized student conversations with Claude and found the dominant use cases were creating and improving academic content — drafting essays, polishing assignments — rather than requesting explanations. The company flagged the pattern itself, noting that direct answer-seeking raises concerns about whether students are building foundational skills. Distinguishing solid AI output from fabricated content is itself a skill many students lack; StudyVerso has covered the basics in how to spot an AI hallucination versus a real fact in 30 seconds.
Delegation or Engagement: Two Ways to Use the Same Tool
Research on AI and learning increasingly converges on one distinction: outcomes depend less on whether students use AI than on how. HEPI’s 2025 data shows explanation-seeking and text-drafting are both mainstream — 42% of students used AI to summarize articles, while roughly one in five pasted AI text directly into assessed work.
The same chatbot can function as a tireless tutor or a ghostwriter. The table below summarizes the two usage patterns and what the available evidence suggests about each.
| Usage pattern | Typical behavior | What the evidence suggests |
|---|---|---|
| Delegation | Generating drafts, solving problem sets, submitting lightly edited AI output | Lower recall and weaker neural engagement (MIT Media Lab, 2025); skills the assessment was meant to build go unpracticed |
| Engagement | Requesting explanations, Socratic questioning, self-testing, critiquing AI output | Consistent with retrieval-practice research; supported by design in OpenAI’s Study Mode and Anthropic’s Learning Mode (2025) |
The AI industry has tacitly conceded the problem by building around it. OpenAI launched Study Mode in July 2025, and Anthropic added Learning Mode to Claude for Education three months earlier; both withhold direct answers and push users through guided reasoning instead. Consumer study apps — from Quizlet to Spanish startups such as Modo Cheto — are converging on similar retrieval-based mechanics. Whether students voluntarily choose the harder mode when the easy one sits one toggle away remains an open empirical question.
What the Potemkin University Trap Means for Students and Institutions
For universities, the risk is credential inflation: degrees that certify less than they claim. For students, it is paying for capability they never acquire. HEPI (2025) found only around a quarter of students felt their institution’s AI guidance was clear and consistently applied, leaving most to navigate the line alone.
Institutional responses are diverging. Some universities are shifting weight toward invigilated exams, oral defenses and in-class writing — assessment formats that AI cannot sit. Others are redesigning assignments to require AI use and grade the student’s judgment about the output instead. Both approaches accept the same premise: the unsupervised essay, as a measure of learning, is losing evidentiary value.
Students face a subtler calculation. Employers are already testing skills directly — live coding rounds, case interviews, work trials — precisely because they trust transcripts less. A student who delegates four years of cognitive work to a model may pass every course and then fail the first unmediated test of what they actually know. Learning to interrogate AI output, rather than accept it, becomes part of the job description; being unable to tell a hallucination from a verifiable fact is now a professional liability, not just an academic one.
The Question Nobody Has Answered Yet
The Potemkin university trap is not inevitable; it is a design problem. The evidence so far suggests AI harms learning when it replaces cognitive effort and can support it when it structures that effort. What no study has yet established is whether institutions can steer millions of students toward the harder path at scale — or whether the façade, cheaper and faster to build than ever, simply becomes the building.