AI Workshop · Charles Darwin University

Cognitive Outsourcing

Handing a piece of thinking to a machine is not new, and it is not always a loss. This section sets out what the research actually shows. The same tool can help a student learn, or can quietly do the learning for them. The difference is in how it is used, not in the tool. It holds the alarming headlines and the enthusiastic ones side by side, and keeps only what the evidence will carry.

Handing work to a tool is old, and usually fine

People have always shifted mental work onto things outside the head. We write a phone number down instead of holding it in memory. We use a calculator instead of doing long division. We set a reminder instead of trusting ourselves to recall. The research term for this is cognitive offloading. A 2016 review defined it as "the use of physical action to alter the information processing requirements of a task so as to reduce cognitive demand". In plain terms, it means letting something outside your head carry part of the load.

That review makes two points worth keeping. Offloading is normal and often useful; it frees attention for other things. It also has a known cost: you tend to remember less of what you hand off. Write the number down and you are less likely to learn it. This trade-off predates generative AI by decades. AI is the newest and most capable thing to offload to, not a new kind of harm.

So the useful question is not whether students use AI. It is what the tool does with the part of the task that the assessment is meant to build. Sometimes the tool supports the student while they do the thinking; that is assistance, or scaffolding. Sometimes the tool does the thinking and the student hands the result in; that is substitution. The whole of the evidence below turns on that distinction.

The distinction that runs through this section

Scaffolding keeps the student doing the work the task is designed to develop, with the tool giving hints, feedback or a way in. Substitution lets the tool supply the finished answer, so the practice that builds the skill never happens. The same product, say ChatGPT, can do either. Which one it does is set by the task and the student, not by the tool.

The clearest evidence: it depends on whether the tool scaffolds or substitutes

The strongest single study is a randomised field experiment with about a thousand high-school mathematics students. It was published in the Proceedings of the National Academy of Sciences in 2025. Students practised with one of three set-ups: no AI, an unrestricted ChatGPT that would give full solutions, or a guardrailed "tutor" version that gave hints but withheld the answer. The researchers then measured two different things, and this is the part that matters.

During practice, with the tool in front of them, both AI groups did better than the no-AI group. The measure that counts came afterwards, on an exam the students sat with no AI at all. There, the picture reversed for one group and held for the other.

During practice, with the tool
+48%  /  +127%
Unrestricted ChatGPT lifted practice performance about 48 per cent over the no-AI group; the hints-only tutor lifted it about 127 per cent.
Later exam, no tool
−17%
The unrestricted-ChatGPT students then scored about 17 per cent worse than the students who had practised with no AI at all.
Later exam, no tool
no gap
The hints-only tutor students were no different from the no-AI group. The guardrail removed the later penalty.

The reason the researchers give is simple. Students with the unrestricted tool tended to ask for and copy the full worked solution. They finished the practice problems without doing the problem-solving, so when the tool was taken away, they had less of the skill the practice was meant to build. The tutor version made them keep working, so it helped in the moment and cost nothing later.

An Australian study points the same way for university writing. Researchers at Monash University ran a randomised experiment with 117 students across four conditions, including ChatGPT and no tool. The ChatGPT group produced the largest jump in essay score. When the researchers measured what the students had actually learned, and whether they could apply it to a new task, the ChatGPT group showed no advantage at all. The authors named the pattern metacognitive laziness. Metacognition is the monitoring you do of your own thinking, noticing what you do not yet understand and deciding what to do about it. When a tool will supply and polish the answer, students tend to stop doing that monitoring themselves and let the tool do it.

Put the two studies together and a single pattern appears. How well the finished work turns out, and how much durable skill the student walks away with, are two separate things. They can move in opposite directions. The gap between them opens where the tool supplies a finished answer and takes away the practice. It stays closed where the tool holds back and makes the student keep working.

SAME TOOL, TWO POSTURES, TWO OUTCOMES higher lower PERFORMANCE During practice (with the tool) Later exam (no tool) No AI (control) hints-only tutor +127% unrestricted AI +48% −17% vs control the practice was done for them Both AI groups looked ahead during practice. Only the group that kept doing the work held the gain.
The pattern drawn from Bastani and colleagues, PNAS 2025 (about 1,000 high-school maths students; the study was run in Turkey). The vertical scale is relative performance, not exam marks; the percentages are the study's reported figures for practice gains and for the later unaided exam. The shape, not the exact height, is the point.
Derek Muller (Veritasium), "What Everyone Gets Wrong About AI and Learning", Perimeter Institute lecture, 2026. A working-through of why effort during practice is the thing that builds durable understanding, and what that means for using AI in study.

The "AI is damaging our brains" claims are weaker than they sound

Alongside the careful studies sit some frightening headlines. A tool is said to be rotting memory, or destroying critical thinking, or leaving a "cognitive debt". These claims travel fast and are worth handling with care, because the evidence under them is thinner than the language suggests.

The most-shared example is an EEG study from the MIT Media Lab titled "Your Brain on ChatGPT". It measured brain activity while people wrote essays with and without a chatbot. Three things about it need stating plainly. It is a preprint, which means it has not yet passed peer review. It is small: 54 people took part, and only 18 completed the session that produced the most-quoted result. And "cognitive debt" is the authors' own turn of phrase, not a measured medical finding; the authors themselves ask readers not to over-interpret it. A formal critique published in late 2025 lists five methodological problems and concludes the study cannot support strong causal claims.

A second example is a 2025 survey of 666 people, backed by 50 interviews, which found that heavier AI users tended to score lower on a critical-thinking measure. The finding is real, but the design is the limit. It is cross-sectional, meaning it takes one snapshot in time. A snapshot can show that two things go together; it cannot show which one causes the other. People who already think less critically may reach for AI more often, rather than AI making them think less. The study can report the association; it cannot settle the direction.

A third, from Microsoft and Carnegie Mellon researchers, surveyed 319 knowledge workers and found that greater confidence in AI went with less self-reported critical thinking. The catch is in the words self-reported: it measures what people said about their effort, not a tested drop in skill, and the people were workers, not students.

What this evidence can and cannot say

There are short-term signals worth watching. There is no sound basis, yet, for saying generative AI causes lasting harm to memory, the brain or the ability to think critically. The studies reaching for that conclusion are short, correlational, self-reported or not yet peer-reviewed, and several carry formal corrections or critiques. Treat the scary version as unproven, and the reassuring version as unproven too.

"Is AI Making College Students Dumber? Ronny Chieng Investigates", The Daily Show, 2025. A comic segment on the same worry; useful for how the "AI is making us stupid" story is told in public, which is the version students meet first.

Even the good news gets retracted

The caution cuts both ways. It is not only the harm studies that turn out to be shaky; the studies showing AI helps can be shaky too. In 2025 a meta-analysis, a study that pools many other studies, reported that ChatGPT improved students' learning, perception and higher-order thinking. It was widely cited. In April 2026 the journal retracted it.

Screenshot of the Nature retraction note for the paper 'The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis' by Jin Wang and Wenxiang Fan. The note, published 22 April 2026, states the editor retracted the paper owing to concerns regarding discrepancies in the meta-analysis that undermine confidence in the validity of the analysis and its conclusions, and that the authors have not responded to correspondence about the retraction.
The retraction note for Wang and Fan, "The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis" (originally Humanities and Social Sciences Communications vol 12, article 621, published 6 May 2025). The editor retracted it on 22 April 2026 over discrepancies in the analysis, and records that the authors did not respond. Source: nature.com. The exhibit is shown as a caution about the evidence base, not as a claim that AI does not help learning.

The point of the exhibit is not that AI fails to help. It is that the research on all of this is young and changes quickly. A single striking result, in either direction, is a weak thing to build a rule on. When a colleague or a headline says the science has settled the question, that is usually the moment to slow down. What holds up is the pattern across the better studies, described above, not any one number.

The other half of the argument: some AI use adds thinking rather than removing it

It would be easy to leave this section as a warning and stop. That would be one-sided, because the same tools can be used in ways that deepen study rather than short-circuit it. The distinction that makes this clear comes from the writer Steven Johnson, who works on Google's NotebookLM. He separates offloading from what he calls uploading. Offloading hands a task away to avoid the effort. Uploading uses the tool to bring in new material, sources, questions, counter-arguments, that the student then has to think through. Same tool, opposite effect on the thinking.

Three habits turn AI use towards the uploading side, and each has a grounding in how learning works.

The first is grounding the tool in real sources. A general chatbot will produce fluent text whether or not it is true. Tools like NotebookLM, or Copilot's Study and Learn agent, can be restricted to documents the student provides: the lecture slides, the set readings, their own notes. Answers then come from that material and cite it, which reduces the tool's tendency to invent. It reduces it, it does not remove it; the detection sections of this workshop explain why no tool can promise the truth of a sentence. Even so, a grounded tool is a different instrument from an open chatbot.

The second is keeping the difficulty in. Learning research has a name for effort that feels hard but pays off later: desirable difficulty. A student can use the tool to make study harder on purpose. They can ask it to quiz them cold before they reopen the notes, or to build spaced practice that returns to a topic just as it starts to fade. This is the opposite of asking for the answer. The tool sets the test; the student still has to sit it.

The third is naming the trap. An instant, polished answer produces a strong feeling of understanding. That feeling is not the same as understanding, and it fades fast; researchers call it the fluency illusion. Knowing the name helps a student notice when a smooth AI answer has given them the sensation of learning without the substance, and go back and do the work.

A companion walkthrough

The interactive deck Cognitive Uploading works through these habits with a hands-on tour of a grounded notebook. It shows how to load your own sources, quiz yourself against them, and use the tool to bring in new angles rather than finished answers. It pairs with this section and with the Copilot Study and Learn session.

"Educating Kids in the Age of A.I.", The Ezra Klein Show, 2025. A conversation that resists both the panic and the hype, and sits with the harder question of what education is for once the tool can produce the finished work.

Holding both, which is the honest position

This section does not resolve into "AI helps" or "AI harms", because the evidence does not. What it supports is narrower and more useful. Assistance that keeps the student doing the assessed thinking tends to help, or at least not to hurt. Assistance that supplies the finished product lifts the immediate result and can leave the durable skill worse. The tool is the same in both cases. What differs is how it is used, which is decided by the task in front of the student and the way it is set.

That is why the practical work in this workshop sits in the next sections rather than this one. If the question is how the tool is used, then the answer is in the design. Set the task and the conditions so that scaffolding is easy and substitution is pointless. The Copilot Study and Learn section takes one tool apart to show the difference in practice. The assessment redesign section asks what a task can still measure once the tool is on every student's account.

A note on TAFE and VET

Almost all of the research above was done in schools and universities. No study was found on AI cognitive offloading in an Australian VET or RTO setting. VET assessment is built around demonstrated skill in real conditions. That is a different thing from an essay or a problem set, so these findings cannot simply be assumed to carry across. That is a gap in the evidence, not a finding that there is no effect. The competency-based side of CDU should read the pattern above as a prompt to watch for, not as a measured result about their students.

Sources

Claims on this page rest on peer-reviewed studies, named preprints handled as preprints, and one retraction record. The main findings were re-checked against the primary source or an open repository record on 23 August 2026. Where a study is correlational, a preprint or commentary, it is labelled as such in the text above. No study located measures effects beyond roughly sixteen weeks, so the long-term picture is genuinely unknown.

  • Cognitive Offloading, Risko and Gilbert, Trends in Cognitive Sciences 20(9), 2016. The definition of cognitive offloading, and the finding that it is normal and useful but carries a cost to unaided recall.
  • Generative AI can harm learning, Bastani and colleagues, Proceedings of the National Academy of Sciences, 2025. The high-school mathematics field experiment; the practice gains and the later unaided-exam penalty for unrestricted AI, with no penalty for the hints-only tutor. A published correction (26 August 2025) fixed an author affiliation only and left the results unchanged.
  • Beware of metacognitive laziness, Fan, Gašević and colleagues, British Journal of Educational Technology, 2025 (Monash University; preprint at arXiv:2412.09315). The 117-student writing experiment: larger essay-score gain with ChatGPT, no gain in knowledge or transfer, and the "metacognitive laziness" framing. The Australian anchor for this section.
  • Your Brain on ChatGPT, Kosmyna and colleagues, arXiv:2506.08872, 2025. An EEG study; a preprint, not peer-reviewed, with 54 participants and heavy drop-out by the most-quoted session. "Cognitive debt" is the authors' framing.
  • Comment on: Your Brain on ChatGPT, Stankovic and colleagues, arXiv:2601.00856, 29 December 2025. A formal critique setting out five methodological concerns with the study above.
  • AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking, Gerlich, Societies 15(1), 2025. The 666-person survey; a cross-sectional design, so it shows association, not cause. A 2025 correction replaced a duplicated table and left the conclusions unchanged.
  • The Impact of Generative AI on Critical Thinking, Lee, Tankelevitch and colleagues, CHI 2025 (Microsoft Research and Carnegie Mellon). A self-report survey of 319 knowledge workers; an association with self-reported effort, not a measured loss of skill, and in workers rather than students.
  • Retraction Note for "The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis", Wang and Fan (original in Humanities and Social Sciences Communications vol 12, article 621, 6 May 2025; retracted 22 April 2026). Used as the caution exhibit above.
  • Cognitive Uploading, Steven Johnson, Adjacent Possible, 2 June 2026. The offloading-versus-uploading distinction and the grounded-notebook habits. Commentary by a practitioner, used as framing.
  • No empirical study on AI cognitive offloading, dependency or skill transfer in an Australian VET or RTO setting was located at the time of checking. The VET position above is recorded as a gap.
Last updated: 23 August 2026