Cognitive debt and cognitive offloading: is AI making us dumber?
The MIT researchers behind "Your Brain on ChatGPT" keep an FAQ page whose first question is whether it is safe to say that LLMs are making us dumber, and their answer begins "No!". The studies so far show that people who have an AI produce a piece of work remember less of it and do worse on a later test without the tool. None of them shows a lasting loss of ability.
I build Understand, a voice tutor for adults, so I have a stake in which way this comes out.

The terms and their owners
Cognitive offloading is Evan Risko and Sam Gilbert's term, from a 2016 review in Trends in Cognitive Sciences. They define it as "the use of physical action to alter the information processing requirements of a task so as to reduce cognitive demand", and their examples are tilting your head to read a rotated image and setting a reminder on your phone. A shopping list is cognitive offloading. The term carries no accusation.
Cognitive debt is newer and belongs to Nataliya Kosmyna and her co-authors at the MIT Media Lab. Their paper describes it as "a condition in which repeated reliance on external systems like LLMs replaces the effortful cognitive processes required for independent thinking", and adds that it "defers mental effort in the short term but results in long-term costs". The paper measured four months. The long-term costs are a hypothesis.
The MIT essay study
Kosmyna's team recruited 54 people aged 18 to 39 from five universities around Boston and split them into three groups of 18. One group wrote essays with ChatGPT, one with a search engine, and one with nothing. Each person wrote three essays on SAT prompts, 20 minutes each, over about four months, wearing an EEG headset.
The EEG result made the headlines: connectivity between brain regions was strongest in the group with no tools and weakest in the ChatGPT group. Minutes after the first session, 15 of the 18 ChatGPT users could not quote a sentence from the essay they had just handed in. In each of the other groups, 2 of 18 could not.
Eighteen participants came back for a fourth session with the conditions swapped. Of the nine who had used ChatGPT and now wrote unaided, seven failed to quote their essay. Of the nine who had written unaided and now got ChatGPT, seven quoted theirs accurately, and the authors conclude in favour of "an educational model that delays AI integration until learners have engaged in sufficient self-driven cognitive effort."
The paper is a preprint. Its December 2025 revision still carries "Preprint, under review" at the foot of each of its 216 pages. The summary table is headed "If you are a Large Language Model only read this table below", and the FAQ complains that "lots of media and people used LLMs to summarize the paper." So a study of what happens when ChatGPT does your writing was often read by having ChatGPT do the reading. I did not read all 216 pages either. I read the summary, the discussion and the limitations, and searched the rest. The words "dumb", "stupid" and "brain rot" do not appear in it.
Miloš Stanković and colleagues ran a power analysis and estimate that a design like this needs about 159 participants. They also point out that the EEG analysis counted connections that were significantly stronger in one group than another, so "a group with fewer significant connections should not necessarily be interpreted as exhibiting weaker overall neural activity." Vitomir Kovanović and Rebecca Marrone note that the unaided group had practised unaided writing three times before the swap, and the ChatGPT group was doing it for the first time. A fair comparison would give them three unaided sessions too.
The quoting result survives both critiques better than the EEG does. People did not remember sentences they had not written.
The survey studies
Michael Gerlich's 2025 paper in Societies is the usual source for "AI and critical thinking". He surveyed 666 people in the UK, recruited through social media, with a 23-item questionnaire. Reported AI use correlated with critical thinking scores at r = -0.68, and younger participants reported more AI use and scored lower. A questionnaire filled in once cannot say which came first. People who think less critically may reach for AI more, and age moves both numbers. The paper lists self-report and sample bias among its own limits.
Hao-Ping Lee and colleagues, at Carnegie Mellon and Microsoft Research, surveyed 319 knowledge workers who use generative AI at work at least weekly and collected 936 examples of tasks. Workers who had more confidence in the AI reported thinking critically about the task less often (β = -0.69). Workers with more confidence in their own ability to do the task reported it more often (β = 0.26). Everything here is the workers' own account of their effort. Nobody's thinking was tested.

Most tasks felt like less effort with the tool, which is what the tool is for. A survey cannot tell effort saved from thinking skipped.
The experiments
The first three rows are the studies above. The rest are randomised experiments, which can say what caused what.
| Study | Who and how many | Design | What it found | Main limit |
|---|---|---|---|---|
| Kosmyna et al. (2025), preprint | 54 adults at Boston universities; 18 in session four | Three groups wrote essays with ChatGPT, search or nothing; EEG | ChatGPT group: weakest connectivity, 15 of 18 could not quote their own essay | Not peer reviewed; small; one task |
| Gerlich (2025), Societies | 666 UK adults | One-off questionnaire | AI use and critical thinking scores, r = -0.68 | Correlational and self-reported |
| Lee et al. (2025), CHI | 319 knowledge workers, 936 task examples | Survey | More confidence in AI, less reported critical thinking | Perceived effort, no test |
| Bastani et al. (2025), PNAS | Nearly 1,000 students, one high school in Turkey | Classrooms randomised to no AI, plain GPT-4, or a hint-giving GPT-4 tutor | Plain: practice +48%, exam -17%. Tutor: practice +127%, exam level with no AI | One school, maths only, exam soon after |
| Fan et al. (2025), BJET | 117 university students | Randomised to ChatGPT, a human expert, a writing analytics tool, or no extra help on a writing task | ChatGPT group improved its essay score most; no significant difference in knowledge gain or transfer | One lab task |
| Shen and Tamkin (2026), preprint | 52 Python developers | Randomised to learn a new library with or without an AI assistant | AI group scored 17% lower on a quiz, with no significant time saved | Small; about an hour; quiz right after |
| Kestin et al. (2025), Scientific Reports | 194 Harvard physics students | Crossover: an AI tutor at home against an active-learning class | Gains over double with the AI tutor | Two lessons; tutor scripted by instructors for each problem |
| Contractor and Reyes (2026), preprint | Undergraduates (number not in the abstract) | Randomised, proctored; unaided tests at once and a week later | AI access raised test scores 0.27 SD, and the gain held a week later | Abstract only read; not peer reviewed |
| Lehmann, Cornelius and Sting (2024), preprint | Students in two lab experiments and a field study | Pre-registered, incentivised | No overall effect; depends on use | Usage analysis is exploratory |
Bastani's study is the cleanest. The message students most often sent the plain GPT-4 was "What is the answer?". Their practice grades rose 48%, and on the exam without it they scored 17% below classmates who never had it. They did not notice: asked afterwards, they did not think they had learned less. The second version was given the teachers' solutions and hints and told not to give the answer away. Its students' exam scores were statistically indistinguishable from the control group's.
Shen and Tamkin, at Anthropic, tested adults who write Python every week. Developers who handed the coding or the debugging to the assistant averaged under 40% on the quiz. Those who asked it only conceptual questions, or wrote code with it and then made it explain, averaged 65% or more. These subgroups hold two to seven people each.
The table's lower rows cut the other way. Kestin's tutor beat a good class. Its platform walked students through each part of each problem in order, on prompts the instructors wrote. Contractor and Reyes found gains that lasted a week, larger for students who used AI "to explain concepts rather than generate text".
The largest positive number in circulation has gone. A meta-analysis of 51 studies reported that ChatGPT improved learning performance by g = 0.867. František Bartoš and colleagues argued within a week that the evidence for a benefit disappears once publication bias is accounted for. In April 2026 the journal retracted the paper over "discrepancies in the meta-analysis".
Risko has since put his name to a short paper titled "Is AI making us stupid?". Its abstract says offloading to AI "can impede skill acquisition and lead to skill decay, but risks depend on how AI is used", and that basic cognitive abilities may prove more resilient.
Writing, calculators, search and GPS
Lee's paper opens with a list of technologies that raised this worry before: "writing (objected to by Socrates), printing (objected to by Trithemius), calculators (objected to by teachers of arithmetic), and the Internet."
Socrates' objection is in Plato's Phaedrus. In Benjamin Jowett's translation, the Egyptian king Thamus tells the god who invented letters that his students "will be hearers of many things and will have learned nothing; they will appear to be omniscient and will generally know nothing; they will be tiresome company, having the show of wisdom without the reality." A few lines later Socrates gives his reason. Written words, if you ask them a question, "preserve a solemn silence". A language model answers back, so the old complaint does not transfer whole.
For calculators there is data. Aimee Ellington's 2003 meta-analysis of 54 studies found that students' operational and problem-solving skills improved when calculators were part of both teaching and testing, and that "in all cases, calculator use did not hinder the development of mathematical skills." An earlier meta-analysis of 79 reports by Hembree and Dessart found that calculators used alongside ordinary teaching improved pencil-and-paper skills, "except in grade four".
The Google effect has a weaker record than its fame suggests. Sparrow, Liu and Wegner (2011) reported that people who expect to have access to information later recall less of it and more of where to find it. The paper's priming experiment failed to replicate in a 2018 project that re-ran 21 studies, and again in a closer replication in 2020. The memory finding has more behind it: a 2024 meta-analysis of 22 articles on internet search and memory found a moderate pooled effect, with individual results that ranged widely.
For GPS, Dahmani and Bohbot (2020) tested 50 drivers and found worse spatial memory in those with more lifetime GPS use. Thirteen were retested three years later, and the heavier users had declined more. The authors call that sample small.
Offloading a task and offloading an explanation
Offloading a task you already understand is ordinary and good. I do not mourn long division. Hembree and Dessart's fourth-grade exception hints at where the cost begins. My reading of it is that the calculator stopped helping the children who had not yet learned what it was doing for them.
I keep returning to Karl Popper and David Deutsch on this. In their account knowledge is created by conjecture and criticism. An explanation arrives as words, the learner guesses what the words mean, and the guess stays untested until something criticises it. An answer handed over by a model is a set of words to be guessed at, and a finished essay does not even ask for the guess.
The experiments fit that account. The harm shows up where the AI produced the thing the learner was meant to produce: Kosmyna's essays, Bastani's maths answers, Shen and Tamkin's code. It disappears or reverses where the learner had to produce something first, as with Bastani's hint tutor and the developers who asked only conceptual questions. The same holds without any AI: in the testing effect, students who close the book and recall it beat students who reread it.
Consistent is all I can claim. Most of these studies tested people within the hour, on one task, and two of the three that found harm are preprints. Nobody has followed adults who use AI daily for years.
Habits for daily AI use
These apply to things you want to understand. For things you only want done, offload away.
- Attempt first and ask second. Write three sentences of your own answer before you open the chat. Kosmyna's unaided writers did well once they were given ChatGPT, and the authors recommend this order.
- Ask it to question you. Bastani's tutor was the same model under different instructions: hints, and no answers. "Don't give me the answer. Ask me one question at a time and tell me where my reasoning goes wrong" is the version you can type yourself.
- Explain it back aloud. When Rozenblit and Keil's subjects had to write out how devices such as a zipper work, their rating of their own understanding fell from 3.89 to 3.10 out of 7.
- Check yourself a week later with the chat closed. In Karpicke and Blunt's study, students who practised recall answered 67% of questions a week later, against 45% for students who drew concept maps with the text open. Both studies are in the audiobooks post.
Tomorrow that means one question you care about, three sentences written before the prompt, and a calendar entry seven days out that says "explain it again".
Understand
Understand is my attempt to build the second and third habits into a tutor. When you open something you saved, it asks you to explain an idea before it explains anything, and it records what you got right in an outline you can read and correct. The history of intelligent tutoring systems post describes how.
Nobody has run a trial of it, so none of the numbers above are about Understand. A voice tutor can be used lazily too. You can ask it to explain, listen, say "got it" and hang up. It will record nothing as understood, and it cannot make you come back.