[nonimo]
EN
Download

De-identified vs pseudonymised: which one is enough for AI?

· Updated · Written and maintained by Joaquín Trapero, Nonimo

De-identified and pseudonymised are not a strong word and a weak word for the same thing. The Privacy Act defines the first and never mentions the second. De-identified data is a result: nobody in it is reasonably identifiable any more, and the Act stops applying to it. Pseudonymised data is a method: the names and numbers were swapped for labels, and a way back was kept.

So the answer to the question most practices are actually asking, whether a pseudonymised client file can go into ChatGPT, is that the method does not decide it. What decides it is what is left in the text, and who ends up reading it.

That is a harder test than it sounds, and the Australian regulator has already applied it to datasets that someone had described as de-identified. Sometimes the label held. Sometimes it did not, and in the clearest failure the difference was a key that could be turned.

De-identified vs pseudonymised: the short answer for Australia

Most of the confusion comes from importing a European vocabulary into an Australian statute. Under the GDPR, pseudonymisation is a defined term with legal consequences. Under the Privacy Act 1988 it is not a term at all, and the only question the Act asks is whether a person is identifiable or reasonably identifiable.

That leaves three things you might have done to a file, and only one of them changes its legal status. The other two change how much you are exposing, which matters, but they leave the file inside the Act.

what you diddefined in the Privacy Act?where the file ends up
Removed or masked the direct identifiersnostill personal information
Pseudonymised, with the key keptnopersonal information for whoever holds the key
De-identified in its release contextyes, section 6(1)outside the Act, for that context

The last row carries a qualifier that most summaries drop: for that context. The same file can be de-identified in one pair of hands and personal information in another, and everything below turns on which pair of hands you are sending it to.

How the UK, Europe and the United States define all three is laid out in our overseas vocabulary guide, and the Australian rules on which tools you may use at all are in the Privacy Act guide.

The Privacy Act defines one word and not the other

Section 6(1) defines personal information as “information or an opinion about an identified individual, or an individual who is reasonably identifiable”. Then it defines the way out: information counts as de-identified when it “is no longer about an identifiable individual or an individual who is reasonably identifiable”.

Read together, those two definitions make de-identified the exact negative of personal information. There is no third category in between for data that is partly protected. Information is either inside the definition or outside it, and the Australian Privacy Principles follow.

Pseudonymity in the APPs is about your customer, not your file

The word that does appear in the Act points somewhere else entirely. APP 2 gives individuals the option of dealing with an organisation anonymously or by pseudonym, unless the law requires identification or it would be impracticable.

The OAIC’s guidance on APP 2 describes a pseudonym as “a name, term or descriptor that is different to an individual’s actual name”, and gives examples like an email address without the person’s name, a forum user name, or an artist’s pen name. It is a right that belongs to the person dealing with you.

It also contains the sentence that settles the vocabulary question for anyone drafting a policy: “The use of a pseudonym does not necessarily mean that an individual cannot be identified.” That was written about customers choosing a nickname. It applies just as well to a practice choosing [PERSON_1].

Why the European word keeps turning up in Australian policies

Many privacy templates and vendor contracts are written with the GDPR in mind, so pseudonymisation arrives in Australian documents already attached to European assumptions. A policy that says client data “will be pseudonymised in line with the Privacy Act” is citing a statute for a word it does not contain.

The practical fix is small. Say which technique you used, and do not claim a legal status the technique cannot reach. The same drafting point comes up in our AI policy template, which was written for a general audience and needs this one sentence adjusted for Australia.

Who holds the key decides which word you can use

The most useful passage in the regulator’s 2018 de-identification guide is an example rather than a definition: a custodian who de-identifies a dataset and keeps a copy of the original.

De-identified vs pseudonymised: where the key sits A client file and its key stay on the practice computer, where the file remains personal information. Only the labelled text travels to the AI provider, which has no key, and whether that text is de-identified is judged in the provider's hands. Your computer client file + the key personal information What leaves [PERSON_1] and the story The AI provider labelled text no key tested in their hands The label hides the name. It does not hide the story.
Same file, two contexts. OAIC, De-identification and the Privacy Act, 21 March 2018

In the OAIC’s words, a retained original “may enable them to re-identify the data subjects”, so “the dataset may be personal information when handled by that custodian, but may be de-identified when handled by a different entity, since the data access environment is different”.

Personal information on your side, possibly not on theirs

Translate that to a practice. You hold the client file and the table that says [PERSON_1] is your client. On your side, the labelled version is personal information, full stop, because you can reverse it in a second. A payslip with a union fee or a leave certificate naming a condition stays sensitive information in your hands, whatever labels went on.

At the AI provider’s end the question is asked again, from scratch. The provider has no table. What it has is the text you sent, the other information it can reach, and the people and systems that can see it. If nobody there could reasonably work out who the client is, the text may be de-identified in that context.

The test applies where the text lands, with a minimum standard

The OAIC makes the recipient’s context the deciding factor in so many words. Entities “will therefore need to consider whether the information would be identifiable or reasonably identifiable in the hands of the other entity. Otherwise, the APPs will apply to the sharing/release of that data.”

It also sets a floor for that assessment. As a minimum, an entity must apply the motivated intruder test: whether a reasonably competent, motivated person with no specialist skills, who has access to the data, the internet and all public documents, could identify someone from it. For data released publicly, the OAIC asks for an assessment “in the round”, against experts too.

That is the test a pseudonymised prompt has to pass. It is also why the method, on its own, never settles the question. The guide on what OpenAI keeps and deletes sets out who, at the provider’s end, can actually see what arrives.

What the OAIC counts as de-identified data: two steps, not one

The OAIC’s guidance describes de-identification as a process with two parts. Step one strips out whatever names the person directly. Step two is one or both of two further moves, points 2 and 3 below, and it is the part most office routines skip.

  1. Direct identifiers out. Name, street address, TFN, Medicare card number, email, mobile.
  2. Everything else that could point back. Rare characteristics, and unusual combinations of ordinary ones, removed or altered.
  3. Controls where the data lands. Who can see it, what they are contractually allowed to do with it, and where it is stored.

The guidance is direct about the trap in step one: “the removal of name, address or other direct identifiers alone may not result in de-identification for the purposes of the Privacy Act.” A practice that stops at step one has done something useful, and has not done the thing the word describes.

Where a label swap sits on that list

Among the techniques the OAIC lists, one describes pseudonymisation almost exactly without using the word. Encryption or hashing of identifiers, it says, are “techniques that will obscure the original identifier, rather than remove it altogether”, usually so that datasets can be linked without sharing identities.

Obscured rather than removed is the whole distinction. A label is step one done reversibly. It is a legitimate tool, and the OAIC treats it as one, but it only ever gets you through the first of the two steps.

The rest of the environment counts too

For the controls in point 3, the OAIC names four parts of a data access environment: other data, people, infrastructure, and governance structures. A research data lab with signed access agreements scores well on all four. A chat window on a consumer account, used by whoever is at the desk, scores poorly on most of them, which is why the same cleaned paragraph can pass in one and fail in the other.

For a disability provider, our NDIS guide takes both steps through an invented progress note.

Two Australian datasets that were called de-identified

In Australia the line between de-identified data and pseudonymised data shows up most clearly in the Commissioner’s own investigations rather than in the guidance, and one of them reached opposite conclusions about different people in the same release.

MBS and PBS, 2016: the provider numbers came back

On 1 August 2016 the Department of Health published on data.gov.au claims data for a 10% sample of people who had claimed Medicare benefits since 1984 or pharmaceutical benefits since 2003. Each patient appeared under an identification number, years of birth stood in for dates of birth, and service dates were shifted by up to 14 days either way.

Medicare provider numbers were handled differently. They were encrypted with a method that, the OAIC later found, was reversible. On 8 September 2016, researchers Chris Culnane, Benjamin Rubinstein and Vanessa Teague of the University of Melbourne told the Australian Bureau of Statistics they had decrypted them. The dataset came down that day, 38 days after it went up.

38 days
From publication on 1 August 2016 to the decryption report on 8 September 2016. OAIC, MBS/PBS data publication

One dataset, two answers

The Commissioner’s findings split along exactly that line. For patients, re-identification would have required unusual features or a very full knowledge of someone’s medical history, and the Commissioner concluded that patients in the dataset were not reasonably identifiable. For them, the release was de-identified.

For providers, a unique identifier “can be derived from the dataset following a process of decryption”. That made their information personal information, and the Commissioner found the Department breached APP 6 by disclosing it, and APPs 1 and 11 in preparing the dataset. An enforceable undertaking was accepted on 23 March 2018. The purpose test behind that APP 6 finding, applied to AI tools, gets its own treatment in the guide to AI under the Privacy Act.

in the 2016 releasehow they were codedcould the code be reversed?what the Commissioner found
Patientsan identification number, year of birth, service dates shifted by up to 14 daysonly with unusual features or a very full knowledge of someone’s medical historynot reasonably identifiable: de-identified
Medicare providersprovider numbers encryptedyes, the encryption was reversiblepersonal information: APP 6 breached

A reversible code is pseudonymisation. When the key, or the method that works as one, is within reach of the recipient, the code is just the identifier in a different font.

Flight Centre, 2017: checked, and still disclosed

The second case is smaller and closer to a working office. In 2017 Flight Centre gave 16 teams at a design jam access to a customer dataset, despite preliminary checks to de-identify or remove personal information. The data included credit card and passport details, and the error was found after it had been available for 36 hours.

The Commissioner determined in December 2020 that Flight Centre interfered with the privacy of almost 7,000 customers, breaching three APPs. The lesson for anyone about to paste a file is not about hackathons. Checking a file and believing it clean are the same experience right up until someone else reads it. When that someone is a chatbot, the first hours after the mistake are the subject of a separate breach guide.

I-MED: what it took for the OAIC to accept the data as de-identified

The case that went the other way is the one most relevant to AI, because it was about training a model. On 19 September 2024 media reports alleged that I-MED Radiology Network had disclosed patient scans to Annalise.ai, a former joint venture with Harrison.ai, to train a diagnostic AI. The OAIC opened preliminary inquiries the next day.

In its report published on 31 July 2025, the Commissioner was satisfied that the patient data had been “de-identified sufficiently that it was no longer personal information for the purposes of the Privacy Act”, and closed the inquiries without further action.

under 30 million
Patient studies I-MED shared with Annalise.ai between April 2020 and January 2022. OAIC report into preliminary inquiries of I-MED, 31 July 2025

What I-MED actually did

The scale is what makes the method worth reading. Between April 2020 and January 2022, I-MED shared under 30 million patient studies and a similar volume of diagnostic reports. The OAIC’s account of what happened before any of it left is specific, and it reads nothing like a label swap.

what I-MED had in placewhat a paste into a chatbot usually has
Two hashing techniques for IDs, names, addresses and phonesa label on the name
Dates shifted to a random date within set yearsthe real dates
Outliers aggregated into large cohortsthe rare detail left in
Text near the edge of each image redactedthe free text untouched
A contract forbidding re-identificationwhatever the account’s terms say
Separate environments and secure storagethe provider’s own systems

Even with all of that, Annalise.ai reported a very small number of occasions where personal information had reached it in error through failures in the process, and that material was deleted or de-identified. The Commissioner still accepted the outcome, because the risk had been reduced to a sufficiently low level and was backed by governance.

The part a busy practice cannot copy

The right column is no criticism of anyone; that is simply what a prompt looks like. A practice cannot shift the dates in a client’s matter without losing the point of asking, and it cannot bind a consumer chatbot to a contract of its own drafting. What those accounts do promise is set out for ChatGPT and for Copilot.

The OAIC was also careful to add that the case study “should not be taken as an endorsement of I-MED’s acts or practices”. It shows that de-identification for AI is achievable in Australia, not that a label on a name achieves it.

Is pseudonymised data enough to put client data into ChatGPT?

For most client work, no. The problem lies in everything the label does not touch: the town, the business, the family member, the dates, the amounts and the story that ties them together. If those still point to one person, the prompt carries personal information and the paste is a disclosure. In litigation the NSW Supreme Court asks a different question, and its practice note judges subpoenaed material by the platform, not by the labels.

The OAIC’s October 2024 guidance on AI products bought off the shelf spells out the consequence with a worked example. Staff at an insurer paste a customer’s claim, health details and all, into a public chatbot to draft an assessment. By doing so, the OAIC says, the insurance company “is disclosing the information to the owners of the chatbot”, and APP 6 applies to that disclosure.

What pseudonymising does buy you

It is not wasted effort, and the regulator says so. The same guidance says the OAIC “expects organisations to consider what information is necessary and assess whether there are ways to minimise the amount of personal information that is input into the AI product”, and suggests techniques that protect privacy while doing it.

Pulling the TFN, the Medicare card number, the contact details and the client’s name out of a prompt is exactly that minimisation. It means a leak would carry less, it removes the identifiers that carry their own legal regimes, and it makes the remaining judgement smaller. That is a genuine gain, and it is the claim worth making. For a clinic the Medicare number is the clearest case, because a principle of its own governs it.

What it does not change

It does not change the legal status of a file that still identifies someone. APP 6 still asks whether the disclosure fits the purpose you collected the information for. APP 8 still applies if the recipient is overseas. And the regulator’s advice on public chatbots is unchanged: it “recommends that organisations do not enter personal information, and particularly sensitive information, into publicly available generative AI tools”.

Whether one particular paste also triggers the notifiable data breach scheme is another matter, with a deadline of its own, worked through in our guide on putting client data into a chatbot. Much depends on the type of account, and our guides to Claude, Gemini and Copilot set out what we found for each.

A client note before and after, on version 0.2.8

Principles are easier to judge with a real example in front of you. The note below is the kind of thing a bookkeeper or a tax agent might want help drafting a reply to. Everything in it is invented, the numbers included: the tax file, Medicare and ABN numbers were generated to pass their checksums.

BEFORE  Client: Marguerite Halloran, TFN 609 905 481, DOB 02/11/1968.
        Medicare 6689 19577 9. Mobile 0491 570 156, m.halloran@example.com.
        14 Wattlebird Cres, Myrtleford VIC 3737. ABN 44 148 175 704.
        She sold the only vet clinic in Myrtleford in March and wants the
        capital gain split with her brother Tobias before the June lodgement.

AFTER   Client: [PERSON_1], TFN [TFN_1], DOB [BIRTH_DATE_1].
        Medicare [MEDICARE_1]. Mobile [PHONE_1], [EMAIL_1].
        [ADDRESS_1], [ADDRESS_2]. ABN [ABN_1].
        She sold the only vet clinic in Myrtleford in March and wants the
        capital gain split with her brother Tobias before the June lodgement.

The AFTER block is the output Nonimo’s engine produced from that note on 23 September 2026, in the 0.2.8 release for both Mac and Windows, with its suggestions accepted and nothing edited by hand.

Look at the last two lines, which are identical before and after. The only vet clinic in one named town, sold in one named month, by a woman with a brother called Tobias. Anyone in Myrtleford could name her. The file is pseudonymised. It is nowhere near de-identified, and no tool could have made it so without deleting the question.

The sentence that decides it is the one you would have to rewrite

The fix is a human edit, and it is usually quick: “a client sold a regional veterinary practice this financial year and wants a capital gain split with a sibling”. The accounting question survives intact. The client does not.

That rewrite is the second step in the OAIC’s process, done by the one person who knows which details matter. The sector version of the same exercise, for progress notes, is in our NDIS guide, and the pages for accountants and law firms show the identifiers each profession handles most.

De-identified vs pseudonymised in your AI policy

The difference between the two words matters most in the one document an auditor, an insurer or a client will actually read. A policy that promises de-identification commits you to a legal result. A policy that describes pseudonymisation commits you to a method, which you can show you followed. The same care applies to the data paragraph of a client engagement letter, where the wrong word becomes a promise to a client.

Name the technique honestly

Say that identifiers are replaced before anything is entered, and that the key stays on the practice's own systems. Do not call the result de-identified.

Name who reads the story

Somebody has to check the remaining text for the details that identify, and that person should be named, not implied.

Keep the key away from the provider

A key uploaded alongside the text turns a pseudonymised prompt back into an identified one.

Name the tools and the account type

A personal login and a company account give different answers.

Keep a record of each decision

The tool, the date, the approver and the reasons.

What an AI policy should say when the practice pseudonymises, in the order an auditor reads it

The security side of this has a date attached. APP 11.3, which applies to information held after 11 December 2024, says the reasonable steps an entity takes to secure personal information include technical and organisational measures. Swapping identifiers out before a prompt is sent counts as a technical measure. The person who reads the story is an organisational one.

If a file is only held for a job that is finished, the Act adds another option you may have forgotten. Under APP 11.2, personal information you have no further permitted use for must be destroyed or de-identified, unless a law requires it to be kept. A file that no longer exists cannot leak. It is also one less thing to weigh when you read what cyber insurance covers.

Where Nonimo fits, and which of the two words it can claim

Nonimo works locally, on the Mac or Windows machine where the file already sits. Before a prompt is sent anywhere, it swaps the identifiers it recognises for labels, lists every swap so a person can reverse any of them, and stores the table linking each label to its value, encrypted, on that same machine.

For Australian practices the recognised set includes the TFN, ABN, ACN and Medicare number, and the Australian home page shows how it works on a practice computer.

That makes it a pseudonymisation tool, and nothing grander. In Privacy Act terms, the note it hands back is still personal information for as long as you hold the table, and deciding what the story gives away remains with whoever knows the client.

What it keeps on the machine is set out on Nonimo’s security page, and the licence terms on the licence page. The guides on using AI with client data are written so that you can weigh tools like ours as well.

Before the paste: one question, asked about the reader

Forget whether the labels went on. The useful question under the Act is about the next reader: holding only the text you plan to send, could they work out who your client is?

A yes, or a shrug, means the text is still personal information and sending it is a disclosure, labels or no labels. The labels were not wasted. They stripped out the numbers that bring their own legal trouble, and shrank what is left to a decision one person can make well. The decision after that, about the tool and the account, begins in the Privacy Act guide.

Sources

Every page below was open in front of us on 23 September 2026.

The client note in this guide is invented, including the town’s only vet clinic. The name, street, email, phone and clinic belong to nobody we know of, and the tax file, Medicare and ABN numbers were generated to pass their checksums.

Common questions

What is the difference between de-identified and pseudonymised data?

De-identified data has a legal status, and pseudonymised data has simply been through a technique. Under section 6(1) of the Privacy Act, information is de-identified when the person is no longer identifiable or reasonably identifiable, and the Act stops applying to it. Pseudonymised means the identifiers were swapped for labels and a way back was kept. The Act does not use that word, and whoever holds the way back is still holding personal information.

Does the Privacy Act define pseudonymisation?

No. The Act defines personal information and de-identified, and nothing in between. The nearest word, pseudonymity, belongs to APP 2, and it means something else: the right of a customer to deal with you under a name that is not theirs. The OAIC's APP 2 guidance adds that using a pseudonym does not necessarily mean the person cannot be identified.

Is pseudonymised data still personal information in Australia?

For the business that holds the key, generally yes. The OAIC's de-identification guidance says a dataset may be personal information when handled by the custodian who kept the original, and de-identified when handled by a different entity. So the same labelled file can be personal information on your desk and, if nothing else in it identifies anyone, something less at the other end.

Can I put pseudonymised client data into ChatGPT?

Only if nobody is reasonably identifiable from what arrives, and a client file rarely clears that bar once the story is left in. If someone still is, the paste is a disclosure under APP 6, and possibly an overseas one under APP 8. The OAIC's best practice advice is to keep personal information out of public generative AI tools altogether, and it expects whatever does go in to be kept to a minimum.

What did the OAIC decide about I-MED and de-identified data?

In a report published on 31 July 2025, the Commissioner was satisfied that patient data I-MED shared with Annalise.ai to train a diagnostic model had been de-identified sufficiently to fall outside the Privacy Act. I-MED had hashed identifiers, shifted dates, aggregated outliers, redacted text near images and bound the recipient by contract. The OAIC said the report is not an endorsement of I-MED's wider practices.

Is encrypting or hashing a client number enough to de-identify it?

No, not by itself. The OAIC lists encryption and hashing as techniques that obscure an identifier rather than remove it. In 2016 the Department of Health released Medicare data with provider numbers encrypted, researchers at the University of Melbourne reversed the encryption, and the Commissioner found the release breached APP 6 for those providers.

What does reasonably identifiable mean under the Privacy Act?

It asks whether identification is both technically possible and reasonably likely, given the information itself, what else the person holding it can reach, and how practical the linking would be. The OAIC sets a minimum test: whether a reasonably competent, motivated person with no specialist skills, internet access and public documents could identify someone. Context decides, so the answer can change with the reader.

Do I have to de-identify client information I no longer need?

Destroy or de-identify, one or the other. APP 11.2 says that when an APP entity no longer needs personal information for any permitted purpose, it must take reasonable steps to destroy it or de-identify it, unless a law requires it to be kept. For a practice, that usually means the retention periods your professional rules and the tax law already set.

Does Nonimo de-identify documents?

No, it pseudonymises them. It replaces identifiers with labels on your computer before the text goes anywhere, shows you what it changed, and keeps the way back encrypted on that same computer. Under the Privacy Act the file stays personal information in your hands. Whether what you send is identifiable at the other end is a judgement about context that remains yours.