Handing work to a tool is old, and usually fine
People have always shifted mental work onto things outside the head. We write a phone number down instead of holding it in memory. We use a calculator instead of doing long division. We set a reminder instead of trusting ourselves to recall. The research term for this is cognitive offloading. A 2016 review defined it as "the use of physical action to alter the information processing requirements of a task so as to reduce cognitive demand". In plain terms, it means letting something outside your head carry part of the load.
That review makes two points worth keeping. Offloading is normal and often useful; it frees attention for other things. It also has a known cost: you tend to remember less of what you hand off. Write the number down and you are less likely to learn it. This trade-off predates generative AI by decades. AI is the newest and most capable thing to offload to, not a new kind of harm.
So the useful question is not whether students use AI. It is what the tool does with the part of the task that the assessment is meant to build. Sometimes the tool supports the student while they do the thinking; that is assistance, or scaffolding. Sometimes the tool does the thinking and the student hands the result in; that is substitution. The whole of the evidence below turns on that distinction.
Scaffolding keeps the student doing the work the task is designed to develop, with the tool giving hints, feedback or a way in. Substitution lets the tool supply the finished answer, so the practice that builds the skill never happens. The same product, say ChatGPT, can do either. Which one it does is set by the task and the student, not by the tool.
The clearest evidence: it depends on whether the tool scaffolds or substitutes
The strongest single study is a randomised field experiment with about a thousand high-school mathematics students. It was published in the Proceedings of the National Academy of Sciences in 2025. Students practised with one of three set-ups: no AI, an unrestricted ChatGPT that would give full solutions, or a guardrailed "tutor" version that gave hints but withheld the answer. The researchers then measured two different things, and this is the part that matters.
During practice, with the tool in front of them, both AI groups did better than the no-AI group. The measure that counts came afterwards, on an exam the students sat with no AI at all. There, the picture reversed for one group and held for the other.
The reason the researchers give is simple. Students with the unrestricted tool tended to ask for and copy the full worked solution. They finished the practice problems without doing the problem-solving, so when the tool was taken away, they had less of the skill the practice was meant to build. The tutor version made them keep working, so it helped in the moment and cost nothing later.
An Australian study points the same way for university writing. Researchers at Monash University ran a randomised experiment with 117 students across four conditions, including ChatGPT and no tool. The ChatGPT group produced the largest jump in essay score. When the researchers measured what the students had actually learned, and whether they could apply it to a new task, the ChatGPT group showed no advantage at all. The authors named the pattern metacognitive laziness. Metacognition is the monitoring you do of your own thinking, noticing what you do not yet understand and deciding what to do about it. When a tool will supply and polish the answer, students tend to stop doing that monitoring themselves and let the tool do it.
Put the two studies together and a single pattern appears. How well the finished work turns out, and how much durable skill the student walks away with, are two separate things. They can move in opposite directions. The gap between them opens where the tool supplies a finished answer and takes away the practice. It stays closed where the tool holds back and makes the student keep working.
The "AI is damaging our brains" claims are weaker than they sound
Alongside the careful studies sit some frightening headlines. A tool is said to be rotting memory, or destroying critical thinking, or leaving a "cognitive debt". These claims travel fast and are worth handling with care, because the evidence under them is thinner than the language suggests.
The most-shared example is an EEG study from the MIT Media Lab titled "Your Brain on ChatGPT". It measured brain activity while people wrote essays with and without a chatbot. Three things about it need stating plainly. It is a preprint, which means it has not yet passed peer review. It is small: 54 people took part, and only 18 completed the session that produced the most-quoted result. And "cognitive debt" is the authors' own turn of phrase, not a measured medical finding; the authors themselves ask readers not to over-interpret it. A formal critique published in late 2025 lists five methodological problems and concludes the study cannot support strong causal claims.
A second example is a 2025 survey of 666 people, backed by 50 interviews, which found that heavier AI users tended to score lower on a critical-thinking measure. The finding is real, but the design is the limit. It is cross-sectional, meaning it takes one snapshot in time. A snapshot can show that two things go together; it cannot show which one causes the other. People who already think less critically may reach for AI more often, rather than AI making them think less. The study can report the association; it cannot settle the direction.
A third, from Microsoft and Carnegie Mellon researchers, surveyed 319 knowledge workers and found that greater confidence in AI went with less self-reported critical thinking. The catch is in the words self-reported: it measures what people said about their effort, not a tested drop in skill, and the people were workers, not students.
There are short-term signals worth watching. There is no sound basis, yet, for saying generative AI causes lasting harm to memory, the brain or the ability to think critically. The studies reaching for that conclusion are short, correlational, self-reported or not yet peer-reviewed, and several carry formal corrections or critiques. Treat the scary version as unproven, and the reassuring version as unproven too.
Even the good news gets retracted
The caution cuts both ways. It is not only the harm studies that turn out to be shaky; the studies showing AI helps can be shaky too. In 2025 a meta-analysis, a study that pools many other studies, reported that ChatGPT improved students' learning, perception and higher-order thinking. It was widely cited. In April 2026 the journal retracted it.
The point of the exhibit is not that AI fails to help. It is that the research on all of this is young and changes quickly. A single striking result, in either direction, is a weak thing to build a rule on. When a colleague or a headline says the science has settled the question, that is usually the moment to slow down. What holds up is the pattern across the better studies, described above, not any one number.
The other half of the argument: some AI use adds thinking rather than removing it
It would be easy to leave this section as a warning and stop. That would be one-sided, because the same tools can be used in ways that deepen study rather than short-circuit it. The distinction that makes this clear comes from the writer Steven Johnson, who works on Google's NotebookLM. He separates offloading from what he calls uploading. Offloading hands a task away to avoid the effort. Uploading uses the tool to bring in new material, sources, questions, counter-arguments, that the student then has to think through. Same tool, opposite effect on the thinking.
Three habits turn AI use towards the uploading side, and each has a grounding in how learning works.
The first is grounding the tool in real sources. A general chatbot will produce fluent text whether or not it is true. Tools like NotebookLM, or Copilot's Study and Learn agent, can be restricted to documents the student provides: the lecture slides, the set readings, their own notes. Answers then come from that material and cite it, which reduces the tool's tendency to invent. It reduces it, it does not remove it; the detection sections of this workshop explain why no tool can promise the truth of a sentence. Even so, a grounded tool is a different instrument from an open chatbot.
The second is keeping the difficulty in. Learning research has a name for effort that feels hard but pays off later: desirable difficulty. A student can use the tool to make study harder on purpose. They can ask it to quiz them cold before they reopen the notes, or to build spaced practice that returns to a topic just as it starts to fade. This is the opposite of asking for the answer. The tool sets the test; the student still has to sit it.
The third is naming the trap. An instant, polished answer produces a strong feeling of understanding. That feeling is not the same as understanding, and it fades fast; researchers call it the fluency illusion. Knowing the name helps a student notice when a smooth AI answer has given them the sensation of learning without the substance, and go back and do the work.
The interactive deck Cognitive Uploading works through these habits with a hands-on tour of a grounded notebook. It shows how to load your own sources, quiz yourself against them, and use the tool to bring in new angles rather than finished answers. It pairs with this section and with the Copilot Study and Learn session.
Holding both, which is the honest position
This section does not resolve into "AI helps" or "AI harms", because the evidence does not. What it supports is narrower and more useful. Assistance that keeps the student doing the assessed thinking tends to help, or at least not to hurt. Assistance that supplies the finished product lifts the immediate result and can leave the durable skill worse. The tool is the same in both cases. What differs is how it is used, which is decided by the task in front of the student and the way it is set.
That is why the practical work in this workshop sits in the next sections rather than this one. If the question is how the tool is used, then the answer is in the design. Set the task and the conditions so that scaffolding is easy and substitution is pointless. The Copilot Study and Learn section takes one tool apart to show the difference in practice. The assessment redesign section asks what a task can still measure once the tool is on every student's account.
Almost all of the research above was done in schools and universities. No study was found on AI cognitive offloading in an Australian VET or RTO setting. VET assessment is built around demonstrated skill in real conditions. That is a different thing from an essay or a problem set, so these findings cannot simply be assumed to carry across. That is a gap in the evidence, not a finding that there is no effect. The competency-based side of CDU should read the pattern above as a prompt to watch for, not as a measured result about their students.
Sources
Claims on this page rest on peer-reviewed studies, named preprints handled as preprints, and one retraction record. The main findings were re-checked against the primary source or an open repository record on 23 August 2026. Where a study is correlational, a preprint or commentary, it is labelled as such in the text above. No study located measures effects beyond roughly sixteen weeks, so the long-term picture is genuinely unknown.
- Cognitive Offloading, Risko and Gilbert, Trends in Cognitive Sciences 20(9), 2016. The definition of cognitive offloading, and the finding that it is normal and useful but carries a cost to unaided recall.
- Generative AI can harm learning, Bastani and colleagues, Proceedings of the National Academy of Sciences, 2025. The high-school mathematics field experiment; the practice gains and the later unaided-exam penalty for unrestricted AI, with no penalty for the hints-only tutor. A published correction (26 August 2025) fixed an author affiliation only and left the results unchanged.
- Beware of metacognitive laziness, Fan, Gašević and colleagues, British Journal of Educational Technology, 2025 (Monash University; preprint at arXiv:2412.09315). The 117-student writing experiment: larger essay-score gain with ChatGPT, no gain in knowledge or transfer, and the "metacognitive laziness" framing. The Australian anchor for this section.
- Your Brain on ChatGPT, Kosmyna and colleagues, arXiv:2506.08872, 2025. An EEG study; a preprint, not peer-reviewed, with 54 participants and heavy drop-out by the most-quoted session. "Cognitive debt" is the authors' framing.
- Comment on: Your Brain on ChatGPT, Stankovic and colleagues, arXiv:2601.00856, 29 December 2025. A formal critique setting out five methodological concerns with the study above.
- AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking, Gerlich, Societies 15(1), 2025. The 666-person survey; a cross-sectional design, so it shows association, not cause. A 2025 correction replaced a duplicated table and left the conclusions unchanged.
- The Impact of Generative AI on Critical Thinking, Lee, Tankelevitch and colleagues, CHI 2025 (Microsoft Research and Carnegie Mellon). A self-report survey of 319 knowledge workers; an association with self-reported effort, not a measured loss of skill, and in workers rather than students.
- Retraction Note for "The effect of ChatGPT on students' learning performance, learning perception, and higher-order thinking: insights from a meta-analysis", Wang and Fan (original in Humanities and Social Sciences Communications vol 12, article 621, 6 May 2025; retracted 22 April 2026). Used as the caution exhibit above.
- Cognitive Uploading, Steven Johnson, Adjacent Possible, 2 June 2026. The offloading-versus-uploading distinction and the grounded-notebook habits. Commentary by a practitioner, used as framing.
- No empirical study on AI cognitive offloading, dependency or skill transfer in an Australian VET or RTO setting was located at the time of checking. The VET position above is recorded as a gap.
