Why this matters
The personal course covered passwords, two-factor authentication and accounts; this page assumes those and focuses on the data you handle as a researcher. The core risk with AI is simple to state: when you type something into a tool, it can leave your control. With many free tools, what you type may be stored and used to help train the next version of the model, which means it could be seen by a reviewer or surface again later.
That is fine for a recipe. It is not fine for information about research participants, communities, or anything held in confidence.
The test that travels everywhere: if you would not put it on a public noticeboard, do not paste it into a public AI tool.
Free, paid and workplace tools are not the same
The single most important choice is which account you use.
- Free public tools generally offer the least protection; assume your inputs may be stored and used for training.
- Paid personal tiers sometimes commit to not training on your inputs; check the specific tool's policy, do not assume.
- Workplace or enterprise accounts, set up by your organisation, are built so your data stays inside the organisation's control and is not used to train the public model.
- Tools that work only from documents you upload (such as NotebookLM) keep the material in one place and reduce invented content, though the same data-sensitivity questions still apply.
For work that touches real information, the safe choice is the tool your organisation has approved, not a personal free account. A worked example is the Northern Territory Government's rollout of Microsoft Copilot: an enterprise deployment is configured so that prompts and data stay within the organisation's tenant and are not used to train the public model. If your workplace provides an approved tool like this, use it for work data, and keep the free public tools for things that are not sensitive.
What never to put into a public AI tool
In a research role, treat all of the following as off-limits for any public or unapproved tool:
- Identifiable or re-identifiable participant or patient information (names with details, addresses, dates of birth, health, financial or other sensitive information).
- Community, cultural, sacred or restricted information, and any community-held data.
- Confidential project, organisational or partner documents.
- Other people's personal details that they have not agreed to share.
- Passwords, access codes or keys.
If an AI tool would genuinely help with a task involving real information, the answer is not to avoid AI; it is to use an approved tool, or to remove the identifying details first.
De-identification, and its limits
De-identification means removing the details that could identify a person before the text goes into a tool. The simple version that always helps: replace specifics with placeholders, "[name]", "[address]", "[DOB]", before you paste. The AI still understands the document and you keep the sensitive parts out of it.
Two cautions matter in remote and community contexts:
- Small numbers re-identify easily. In a small community, "a 54-year-old man from [town] with [condition]" can identify someone even with the name removed. De-identification is about whether a person could be picked out, not just whether their name is present.
- AI can infer identity. The Office of the Australian Information Commissioner has flagged that an AI inferring personal details can itself amount to collecting personal information. Stripping the obvious fields does not guarantee the data is safe to enter.
When in doubt, do not paste it; ask first.
The rules that apply
You do not need to be a lawyer, but it helps to know the shape of the rules.
Privacy Act 1988 (Commonwealth), with the 2024 reforms. Health information is "sensitive information" with heightened protection, and health-sector bodies are covered regardless of size. From late 2026, organisations must be transparent in their privacy policies about automated decision-making, defined broadly enough to include AI. Importantly, the NHMRC guidelines that let researchers use some health data without individual consent apply to the data itself; they do not authorise uploading identifiable or re-identifiable health data to a commercial AI platform. And under the reformed cross-border rules, no countries have yet been approved for routine overseas data transfer, which matters because most AI tools process data overseas.
NT Information Act 2002. The Northern Territory does not have a privacy act mirroring the Commonwealth one; instead the Information Act sets Information Privacy Principles for NT Government agencies. If you work with NT Government health data, both Commonwealth and Territory requirements can apply.
NHMRC expectations. Using AI to search, summarise or code literature is treated as routine research use, while using AI to process participant data or make predictions triggers fuller ethical assessment. Grant applicants may use AI but must certify accuracy; peer reviewers must not use AI to assess applications.
The detail of these changes over time. The point to carry is the principle: identifiable and sensitive data needs an approved, controlled tool and a lawful basis, not a free public chatbot.
The ICIP thread runs through all of this
Responsible data handling and Indigenous Data Sovereignty are the same conversation here. Community, cultural and identifying data carries obligations that go beyond privacy law, to the authority of Aboriginal and Torres Strait Islander people over their own data and knowledge. Putting that data into a public AI tool is not only a privacy risk; it can cut across data sovereignty and consent. The frameworks, and the questions worth asking, are on the data sovereignty and ICIP page; where a question touches cultural material or community authority, it goes to the relevant Aboriginal and Torres Strait Islander leadership and governance bodies, not to a tool.
A practical checklist
- Use your organisation's approved AI tool for anything work-related; keep free tools for non-sensitive tasks.
- Before pasting, ask: could a person be identified from this, directly or by inference, especially in a small community?
- De-identify with placeholders, and remember small numbers and inference can still re-identify.
- Never put community, cultural, sacred or restricted information into a public tool.
- Check whether the work needs a lawful basis (consent, NHMRC guidelines, ethics approval) before any data goes near a tool.
- If you are unsure, ask before you paste; ask your data custodian, your supervisor, or the relevant governance body.
Glossary
- Personal information
- information that can identify a person, on its own or combined with other information.
- Sensitive information
- a protected category under the Privacy Act that includes health information.
- De-identification
- removing details so a person cannot reasonably be identified; harder in small communities.
- Re-identification
- working out who someone is from supposedly de-identified data.
- Enterprise / workplace tool
- an AI account configured by an organisation so data stays in its control and is not used to train the public model.
- Tenant
- an organisation's own walled-off space in a cloud service, where its data is kept separate.
- Privacy Act 1988
- the main Commonwealth privacy law; reformed in 2024.
- NT Information Act 2002
- the Northern Territory law setting Information Privacy Principles for NT Government agencies.
Take it with you: a one-page "what to put in, what to keep out" cheat sheet is available to download from the button at the top of this page.
