AI Workshop · Charles Darwin University

Assessment Redesign

The aim of a redesign is not a task that AI cannot touch. The aim is a task whose result still supports a defensible claim about what a student can do. This section sets out what the Australian regulators actually ask, what each way of assessing can and cannot establish, and how to review a task you are uneasy about. Higher education and VET are treated separately, because they start from different rules.

The question is validity, not whether the task is AI-proof

Validity is the property that matters here. An assessment is valid when the inference it licenses holds: the inference from the work a student submits to a claim about what that student can do. Generative AI is a threat to that inference, not only a cheating problem. If a tool can produce the part of the work that was meant to show the capability, the result no longer evidences the capability.

Australian scholarship has made this the working frame. Dawson, Bearman, Dollinger and Boud argue that validity matters more than cheating. Their point is that catching misuse is secondary; the first question is whether a result still means what the institution says it means. A task can be made harder to complete with AI and still fail this test. Making a task AI-resistant is not the same as obtaining valid evidence of learning.

Take a plain example. A unit sets a take-home essay that asks the student to interpret a supplied data set and argue from it in their own words. The capability claimed is independent analysis of data. If a current tool can read the brief, take the data set and return a passing argument, then a clean submission no longer tells you the student can analyse data. The task still looks rigorous. The inference underneath it has quietly broken.

One consequence is uncomfortable and worth stating early. Securing a task can raise the fidelity of the evidence while narrowing what the task measures. A supervised, closed-room version of that essay assures you the student wrote it. It may also reduce the task to what a person can produce in an hour without resources, which is a smaller thing than the original construct. Security and validity are distinct, and buying one can cost some of the other.

The frame for this whole section

Ask of any assessment: does the result still support the claim you want to make about the student, now that AI is available? Redesign is whatever makes the honest answer yes. It is rarely the same as making the task impossible to complete with a tool.

What TEQSA asks of higher education

TEQSA is the higher education regulator. Its September 2025 report, Enacting assessment reform in a time of artificial intelligence, is the current anchor for the sector. It shifts the emphasis away from detection and onto the redesign of assessment, so that tasks capture authentic demonstrations of what a student can do. Two guiding principles sit at the top of that report.

The first principle is that assessment and learning experiences "should equip students to participate ethically, critically and actively in a society where gen AI is ubiquitous". The second is that "forming trustworthy judgements about student learning requires multiple, inclusive and contextualised approaches to assessment". The word trustworthy is doing the work. Trust is built across a program from several points of evidence, not extracted from one policed task.

For any single task, the same report offers three options rather than a demand to make everything exam-like. AI use can be permitted within defined parameters. The task can be designed so that AI use is irrelevant to what it measures. Or use can be restricted through direct supervision. Each option is a legitimate response; which one fits depends on what the task is for. TEQSA's own term for a supervised or in-person checkpoint is a "secure assessment task", such as an interactive oral or an in-class demonstration.

The regulator also warns against the reflex that AI anxiety produces. Under pressure, institutions may "revert to inequitable assessment formats". They opt for "familiar, secure assessment formats such as invigilated time-limited tests and exams". These, the report says, may "reduce assessment variety and authenticity" and exacerbate "existing inequities". This is the regulator's own caution, not commentary. Defaulting every task to a supervised exam is a validity and equity cost, taken to buy security.

Where to place the secured points

If trustworthy judgements are built across a program, the design question becomes where the secured points sit. A secured point is a task where you can be confident the work is the student's own, because it is supervised, oral or in-class. The TEQSA report sets out three ways to place these points, and each carries a workload trade-off.

WHERE THE SECURED POINTS SIT ACROSS A PROGRAM PROGRAM-WIDE reform Few secured points, at progression moments. Least total workload; needs program coordination. UNIT-LEVEL assurance A secured point in every unit. Quick to adopt; heavy scheduling and supervision overhead. HYBRID model Secured points on the units that matter most; program principles applied to the rest. start of program completion secured point open or AI-integrated unit
The three placements TEQSA sets out in its September 2025 report, drawn for this course over a program of eight units. The point is not that one placement is correct. It is that the number and position of secured points is a design choice with a workload cost, made at program level rather than task by task.

Program-wide reform treats the whole course as one system of evidence. A small number of secured points sit at moments that matter for progression and completion, and the rest of the assessment can be open and AI-integrated. Total workload is lower, because supervision is not duplicated in every unit. The cost is coordination: someone has to hold the program view and decide where the secured points belong.

Unit-level assurance puts a secured point in every unit. It is quick to adopt, because each unit is fixed on its own without a program conversation. The cost lands later, in scheduling, supervision and documentation repeated across every unit at once. Hybrid designs sit between the two. They secure the units that carry the most weight and apply program principles elsewhere. No Australian study yet measures which placement best assures learning; the report endorses the program view on reasoning, not on outcome data.

The two-lane approach, and the argument about it

One design has become common enough to name. It is usually called the two-lane approach, and the language of "secured" and "open" lanes comes from the assessment literature, not from the regulator. A secured lane verifies that a student has attained the outcomes independently. An open lane develops and assesses the productive use of AI. The University of Sydney and the Australian Catholic University both publish worked versions. The stated logic is that neither lane alone is enough: security without open tasks trains nobody to use the tools, and open tasks alone assure nothing.

The approach is useful, and it is also contested in the peer-reviewed literature. A University of Western Australia critique carries the title "The two-lane road to hell is paved with good intentions". It argues that an all-or-none split between secured and open tasks is hard to sustain, and raises coherence and equity concerns. A separate Australian paper, "Talk is cheap", makes a related point. Surface-level fixes, such as asking students to use AI appropriately or tweaking a prompt, do not restore validity. Structural change across a program does. Present the two-lane model as a leading Australian option under active debate, not as a settled solution. What the evidence supports is the direction, program-level assurance with secured points, more than any single packaged model.

What each way of assessing can and cannot establish

Redesign eventually comes down to choosing a mode of assessment. Every mode buys some assurance and leaves something unestablished, and most trade assurance against staff workload. The table sets out what the located Australian evidence supports for the common modes. One line runs through all of them: no single mode establishes independent achievement on its own, which is the reason the frameworks reach for combinations and program-level triangulation.

ModeWhat it can establishWhat it cannotWorkload and equity
Open, AI-permitted task Whether a student can work productively and critically with AI, which is itself a capability worth assessing. That the underlying work is the student's own, on its own; a clean result does not show independent achievement. Scales well. Needs parameters written against the outcome, or it measures nothing in particular.
Oral or interactive oral Whether a student can explain and defend the work in real time, which probes understanding a script cannot fake easily. A measured reduction in cheating; the Australian evidence shows students perceive it as harder to cheat, not that incidence falls. Scales poorly; time-intensive. Raises language, anxiety and reasonable-adjustment questions for EAL and disabled students.
Practical demonstration or observation Real-time performance to a standard, watched by the assessor as the person performs it. Underpinning reasoning, which needs paired questioning, or consistency, which needs a range of situations. One-to-one or small-group and assessor-time-intensive. Native to VET competency-based assessment.
Workplace or work-integrated evidence Performance in a real setting, which carries the highest ecological validity of any mode. Assurance the RTO can stand behind directly; the authenticity risk shifts to a third-party collector. Placement access is itself an equity issue for remote, regional and carer students; supervisor variability is a known weak point.
Proctored online exam A deterrent effect, on the accounts of Australian staff who have used the software. Reliable identity checking or cheating detection; staff described the tooling as unreliable and used it as deterrence. Documented accessibility harms: facial recognition and darker skin, eye-tracking and autistic students, ID checks and trans or undocumented students.
Program-level verification A capability across several points of evidence, so that no single task carries the whole authenticity load. A guarantee at scale; the model is endorsed in principle and not yet empirically validated in the located Australian evidence. Lower duplicated workload than securing every unit, at the cost of program coordination.

Two practical readings follow. First, the mode has to be chosen against the capability being claimed, not against how AI-resistant it feels. An oral defence is powerful for reasoning and weak for a portfolio of applied work. Second, the equity effects are mode-specific, and some are documented rather than speculative. A redesign that fixes a validity problem by creating an access problem has not improved the assessment; it has moved the fault.

VET starts from a different rule: authenticity

VET does not approach this through validity theory. It starts from the rules of evidence under the Standards for Registered Training Organisations 2025, the framework the VET regulator ASQA administers. Assessment evidence must satisfy four rules: it must be valid, sufficient, authentic and current. Authenticity is the rule that AI presses on. The ASQA practice guide defines it as the assessor being "assured that a VET student's assessment evidence is the original and genuine work of that VET student".

That wording changes the shape of the problem. In higher education, an authenticity concern usually surfaces as a contested misconduct allegation against a student. In VET, authenticity is an affirmative assurance the assessor has to reach before recording competence at all. If the assessor cannot be assured, the outcome is not yet competent, or not able to be assessed on this evidence. It is a competency-judgement outcome, not automatically a misconduct finding. The guide names AI directly. It asks assessors to validate that evidence "has not been plagiarised or generated with artificial intelligence (AI) tools". It also asks them to verify that the person enrolled, trained and assessed is the same person.

This gives VET assessors tools that the higher education essay lacks. Competency-based assessment already uses live observation, oral questioning, contextualised practical demonstration and identity verification. Where those are used, the authenticity assurance is reached through the method itself, not through a detector run after the fact. CDU's own VET Assessment System Policy localises the ASQA structure and requires assessors to verify that evidence is the student's own work. CDU's Generative AI Policy adds the hard rule on the marking side. A VET competency judgement must be made by a qualified assessor, and an AI platform is not one.

TAFETalks: Navigating Assessment and Integrity in the Age of GenAI. A VET-sector panel on how the rules of evidence apply when candidates can use generative AI. Listed on the resources page with the rest of the recommended viewing.

The gap where student-side use should be

ASQA has developed principles for the responsible use of AI in VET. They took shape in draft through the first half of 2026 and were released around the middle of the year. They matter for what they cover, and for what they leave to the existing rules. The principles are addressed to providers and RTOs: governance of AI, human oversight and accountability, secure handling of information, student equity and wellbeing, and alignment to the training product. They tie into the current Standards rather than adding new legislation.

What they do not do is set rules for how a candidate may use AI in producing assessment evidence. That is the gap. ASQA's AI material governs how a provider deploys AI, including how an RTO uses it near students, but not the candidate-side question this workshop keeps returning to. In VET, that question is not left unanswered; it falls back on the authenticity rule of evidence. The assessor still has to be assured the work is the candidate's own, whatever tools were available. So VET does have a clear standard for candidate AI use. It sits in the rules of evidence, not in a dedicated AI instrument, and a reader looking for the latter finds the space where it would go.

For VET participants

You already hold the tool the higher education sector is reaching for. Authenticity is a named rule you assess against, and observation and oral questioning let you reach it directly. The redesign question in VET is often less "how do I detect AI" and more "does this task give me a way to be assured, without a detector".

Reviewing the task you brought

This is the point of the section: a task you actually set, looked at through the frame above. The questions below are the review lens. They work well in pairs. You swap a task with a colleague and review each other's, because a weak point is easier to see in someone else's work than in your own. The materials are the questions and a worked example; run them on the task you brought.

Select a question to see what a strong answer looks like.

What the task is for
Write the claim as a sentence: "a student who passes this can ...". If you cannot, the task has no clear construct to protect, and no redesign will fix that first. The claim is the thing AI availability either does or does not undermine. Everything else in the review is measured against it.
In a writing unit, the writing is the construct. In a chemistry report, the writing is a vehicle and the chemistry is the construct. This decides which kinds of AI help are assistance and which are substitution. The same use of a tool can be either, depending on this answer.
How exposed it is
Try it, honestly, with a current model. If a clean pass comes back from the brief and any supplied materials, the inference from a good submission to the claimed capability is already weak. This is not proof a student did so; it is evidence the task can no longer carry the claim by itself.
Map the task against the range of AI use set out in the declarations section, from spelling help to producing the assessed reasoning. Mark where, for this task, use stops being assistance and becomes substitution. That point is set by the construct, not by the tool, and it is different for a report and for an essay.
What can change
This is the first of TEQSA's three options. Name the uses that are in and out for this task, tied to the construct, and assess the student's judgement in using the tool. It keeps the task open and authentic. It works only where working-with-AI is part of, or compatible with, the capability you are claiming.
The second option. Shift what is assessed onto something a tool does not supply: the student's own data, a local or personal context, an in-process artefact, a defence of choices. The aim is a task where using AI neither helps nor hurts the thing being measured, so the question of misuse partly dissolves.
The third option, and the one to use sparingly. A secured point assures the work is the student's own, and narrows what the task can measure to what a person produces under supervision. Reserve it for the capability that genuinely needs the assurance, and remember the regulator's warning about defaulting to invigilated exams.
How it will be assured
If a later secured task already assures the same capability, this task may not need to carry the assurance itself, and can stay open. This is the program view in practice. It is also why the redesign of one task is hard to finish without a glance at the units around it.
If the task allows live observation, a practical demonstration or a few oral questions, the assessor can reach the authenticity assurance through the method, with no detector involved. Where it does not, add a short questioning component rather than relying on the submitted text alone. The judgement stays with a qualified assessor.
Worked example: the review applied

A 2,000-word take-home essay analysing a supplied workplace case study, submitted online, worth 40 per cent of a management unit.

The claim
A student who passes can analyse an organisational situation and argue a recommendation from evidence. The writing is a vehicle; the analysis is the construct.
Exposure
A current tool, given the case study and the brief, returns a competent essay with a defensible recommendation. The task no longer evidences the claim on its own.
Option taken
Redesign so AI use is irrelevant to the claim. The case study is replaced with one the student sources from their own workplace or placement, and the brief asks for a recommendation defended against two named alternatives specific to that setting.
Assurance
A ten-minute oral follows, where the student explains the local constraints and answers two questions about their recommendation. The oral is the secured point; the essay stays open and AI use is declared, not banned.
What it costs
The oral adds assessor time, and sourcing a local case demands more of the student. In return, the result supports the original claim again, and a tool cannot supply the part that matters.
What the review does not promise

A reviewed task is defensible, not tamper-proof. A determined student can still misuse a tool, and no redesign changes that. What the review buys is a task whose result still means what you need it to mean, and a clear account of how that meaning is assured. That is the achievable goal, and it is the one the evidence supports.

Not AI-proof, but defensible

The through-line of this section is a single shift. Stop asking whether a task can be beaten by a tool, and start asking whether its result still supports the claim you make about a student. That is validity, and it is the question both regulators have settled on, in their different vocabularies. Higher education reaches it through assessment redesign and program-level assurance. VET reaches it through the rules of evidence, with authenticity as a named obligation on the assessor.

The practical work is modest and specific. Name the capability a task claims. Test how exposed it is. Choose, at program level where you can, whether to permit AI within parameters, design it out, or secure a point. Then say how the result is assured. A task handled this way is not immune to misuse. It is defensible, which is the standard an integrity process and an external reviewer will actually hold it to. The declarations and detection sections cover what to do when a concern still arises; this section is about the task that gives them less to do.

Sources

Claims on this page rest on Australian regulator guidance and peer-reviewed research, with institutional practice and commentary labelled where used. Quotations were checked against the primary text or an open repository copy in August 2026. The ASQA AI principles are described from ASQA's public summaries and dated sector reporting, because the primary pages could not be read directly at the time of checking. Their exact wording and release date are on the verification list, not asserted here.

Last updated: 25 August 2026