ntworld.ink
AI for Research · AI in Research Practice

Handling Data Responsibly with AI Tools

What you can and cannot put into an AI tool in a research role, the difference between free and workplace tools, how to de-identify, and the rules that apply to research data.

// take it with you Download cheat sheet (Word)
// on this page
// why this matters

Why this matters

The personal course covered passwords, two-factor authentication and accounts; this page assumes those and focuses on the data you handle as a researcher. The core risk with AI is simple to state: when you type something into a tool, it can leave your control. With many free tools, what you type may be stored and used to help train the next version of the model, which means it could be seen by a reviewer or surface again later.

That is fine for a recipe. It is not fine for information about research participants, communities, or anything held in confidence.

The test that travels everywhere: if you would not put it on a public noticeboard, do not paste it into a public AI tool.

// free, paid and workplace tools are not the same

Free, paid and workplace tools are not the same

The single most important choice is which account you use.

For work that touches real information, the safe choice is the tool your organisation has approved, not a personal free account. A worked example is the Northern Territory Government's rollout of Microsoft Copilot: an enterprise deployment is configured so that prompts and data stay within the organisation's tenant and are not used to train the public model. If your workplace provides an approved tool like this, use it for work data, and keep the free public tools for things that are not sensitive.

// what never to put into a public ai tool

What never to put into a public AI tool

In a research role, treat all of the following as off-limits for any public or unapproved tool:

If an AI tool would genuinely help with a task involving real information, the answer is not to avoid AI; it is to use an approved tool, or to remove the identifying details first.

// de-identification, and its limits

De-identification, and its limits

De-identification means removing the details that could identify a person before the text goes into a tool. The simple version that always helps: replace specifics with placeholders, "[name]", "[address]", "[DOB]", before you paste. The AI still understands the document and you keep the sensitive parts out of it.

Two cautions matter in remote and community contexts:

When in doubt, do not paste it; ask first.

// the rules that apply

The rules that apply

You do not need to be a lawyer, but it helps to know the shape of the rules.

Privacy Act 1988 (Commonwealth), with the 2024 reforms. Health information is "sensitive information" with heightened protection, and health-sector bodies are covered regardless of size. From late 2026, organisations must be transparent in their privacy policies about automated decision-making, defined broadly enough to include AI. Importantly, the NHMRC guidelines that let researchers use some health data without individual consent apply to the data itself; they do not authorise uploading identifiable or re-identifiable health data to a commercial AI platform. And under the reformed cross-border rules, no countries have yet been approved for routine overseas data transfer, which matters because most AI tools process data overseas.

NT Information Act 2002. The Northern Territory does not have a privacy act mirroring the Commonwealth one; instead the Information Act sets Information Privacy Principles for NT Government agencies. If you work with NT Government health data, both Commonwealth and Territory requirements can apply.

NHMRC expectations. Using AI to search, summarise or code literature is treated as routine research use, while using AI to process participant data or make predictions triggers fuller ethical assessment. Grant applicants may use AI but must certify accuracy; peer reviewers must not use AI to assess applications.

The detail of these changes over time. The point to carry is the principle: identifiable and sensitive data needs an approved, controlled tool and a lawful basis, not a free public chatbot.

// the icip thread runs through all of this

The ICIP thread runs through all of this

Responsible data handling and Indigenous Data Sovereignty are the same conversation here. Community, cultural and identifying data carries obligations that go beyond privacy law, to the authority of Aboriginal and Torres Strait Islander people over their own data and knowledge. Putting that data into a public AI tool is not only a privacy risk; it can cut across data sovereignty and consent. The frameworks, and the questions worth asking, are on the data sovereignty and ICIP page; where a question touches cultural material or community authority, it goes to the relevant Aboriginal and Torres Strait Islander leadership and governance bodies, not to a tool.

// a practical checklist

A practical checklist

// glossary

Glossary

Personal information
information that can identify a person, on its own or combined with other information.
Sensitive information
a protected category under the Privacy Act that includes health information.
De-identification
removing details so a person cannot reasonably be identified; harder in small communities.
Re-identification
working out who someone is from supposedly de-identified data.
Enterprise / workplace tool
an AI account configured by an organisation so data stays in its control and is not used to train the public model.
Tenant
an organisation's own walled-off space in a cloud service, where its data is kept separate.
Privacy Act 1988
the main Commonwealth privacy law; reformed in 2024.
NT Information Act 2002
the Northern Territory law setting Information Privacy Principles for NT Government agencies.

Take it with you: a one-page "what to put in, what to keep out" cheat sheet is available to download from the button at the top of this page.