De-identified vs pseudonymised: which one is enough for AI?
· Updated · Written and maintained by Joaquín Trapero, Nonimo
De-identified and pseudonymised are not a strong word and a weak word for the same thing. The Privacy Act defines the first and never mentions the second. De-identified data is a result: nobody in it is reasonably identifiable any more, and the Act stops applying to it. Pseudonymised data is a method: the names and numbers were swapped for labels, and a way back was kept.
So the answer to the question most practices are actually asking, whether a pseudonymised client file can go into ChatGPT, is that the method does not decide it. What decides it is what is left in the text, and who ends up reading it.
That is a harder test than it sounds, and the Australian regulator has already applied it to datasets that someone had described as de-identified. Sometimes the label held. Sometimes it did not, and in the clearest failure the difference was a key that could be turned.
De-identified vs pseudonymised: the short answer for Australia
Most of the confusion comes from importing a European vocabulary into an Australian statute. Under the GDPR, pseudonymisation is a defined term with legal consequences. Under the Privacy Act 1988 it is not a term at all, and the only question the Act asks is whether a person is identifiable or reasonably identifiable.
That leaves three things you might have done to a file, and only one of them changes its legal status. The other two change how much you are exposing, which matters, but they leave the file inside the Act.
| what you did | defined in the Privacy Act? | where the file ends up |
|---|---|---|
| Removed or masked the direct identifiers | no | still personal information |
| Pseudonymised, with the key kept | no | personal information for whoever holds the key |
| De-identified in its release context | yes, section 6(1) | outside the Act, for that context |
The last row carries a qualifier that most summaries drop: for that context. The same file can be de-identified in one pair of hands and personal information in another, and everything below turns on which pair of hands you are sending it to.
How the UK, Europe and the United States define all three is laid out in our overseas vocabulary guide, and the Australian rules on which tools you may use at all are in the Privacy Act guide.
The Privacy Act defines one word and not the other
Section 6(1) defines personal information as “information or an opinion about an identified individual, or an individual who is reasonably identifiable”. Then it defines the way out: information counts as de-identified when it “is no longer about an identifiable individual or an individual who is reasonably identifiable”.
Read together, those two definitions make de-identified the exact negative of personal information. There is no third category in between for data that is partly protected. Information is either inside the definition or outside it, and the Australian Privacy Principles follow.
Pseudonymity in the APPs is about your customer, not your file
The word that does appear in the Act points somewhere else entirely. APP 2 gives individuals the option of dealing with an organisation anonymously or by pseudonym, unless the law requires identification or it would be impracticable.
The OAIC’s guidance on APP 2 describes a pseudonym as “a name, term or descriptor that is different to an individual’s actual name”, and gives examples like an email address without the person’s name, a forum user name, or an artist’s pen name. It is a right that belongs to the person dealing with you.
It also contains the sentence that settles the vocabulary question for anyone drafting a policy: “The use of a pseudonym does not necessarily mean that an individual cannot be identified.” That was written about customers choosing a nickname. It applies just as well to a practice choosing [PERSON_1].
Why the European word keeps turning up in Australian policies
Many privacy templates and vendor contracts are written with the GDPR in mind, so pseudonymisation arrives in Australian documents already attached to European assumptions. A policy that says client data “will be pseudonymised in line with the Privacy Act” is citing a statute for a word it does not contain.
The practical fix is small. Say which technique you used, and do not claim a legal status the technique cannot reach. The same drafting point comes up in our AI policy template, which was written for a general audience and needs this one sentence adjusted for Australia.
Who holds the key decides which word you can use
The most useful passage in the regulator’s 2018 de-identification guide is an example rather than a definition: a custodian who de-identifies a dataset and keeps a copy of the original.
In the OAIC’s words, a retained original “may enable them to re-identify the data subjects”, so “the dataset may be personal information when handled by that custodian, but may be de-identified when handled by a different entity, since the data access environment is different”.
Personal information on your side, possibly not on theirs
Translate that to a practice. You hold the client file and the table that says [PERSON_1] is your client. On your side, the labelled version is personal information, full stop, because you can reverse it in a second. A payslip with a union fee or a leave certificate naming a condition stays sensitive information in your hands, whatever labels went on.
At the AI provider’s end the question is asked again, from scratch. The provider has no table. What it has is the text you sent, the other information it can reach, and the people and systems that can see it. If nobody there could reasonably work out who the client is, the text may be de-identified in that context.
The test applies where the text lands, with a minimum standard
The OAIC makes the recipient’s context the deciding factor in so many words. Entities “will therefore need to consider whether the information would be identifiable or reasonably identifiable in the hands of the other entity. Otherwise, the APPs will apply to the sharing/release of that data.”
It also sets a floor for that assessment. As a minimum, an entity must apply the motivated intruder test: whether a reasonably competent, motivated person with no specialist skills, who has access to the data, the internet and all public documents, could identify someone from it. For data released publicly, the OAIC asks for an assessment “in the round”, against experts too.
That is the test a pseudonymised prompt has to pass. It is also why the method, on its own, never settles the question. The guide on what OpenAI keeps and deletes sets out who, at the provider’s end, can actually see what arrives.
What the OAIC counts as de-identified data: two steps, not one
The OAIC’s guidance describes de-identification as a process with two parts. Step one strips out whatever names the person directly. Step two is one or both of two further moves, points 2 and 3 below, and it is the part most office routines skip.
- Direct identifiers out. Name, street address, TFN, Medicare card number, email, mobile.
- Everything else that could point back. Rare characteristics, and unusual combinations of ordinary ones, removed or altered.
- Controls where the data lands. Who can see it, what they are contractually allowed to do with it, and where it is stored.
The guidance is direct about the trap in step one: “the removal of name, address or other direct identifiers alone may not result in de-identification for the purposes of the Privacy Act.” A practice that stops at step one has done something useful, and has not done the thing the word describes.
Where a label swap sits on that list
Among the techniques the OAIC lists, one describes pseudonymisation almost exactly without using the word. Encryption or hashing of identifiers, it says, are “techniques that will obscure the original identifier, rather than remove it altogether”, usually so that datasets can be linked without sharing identities.
Obscured rather than removed is the whole distinction. A label is step one done reversibly. It is a legitimate tool, and the OAIC treats it as one, but it only ever gets you through the first of the two steps.
The rest of the environment counts too
For the controls in point 3, the OAIC names four parts of a data access environment: other data, people, infrastructure, and governance structures. A research data lab with signed access agreements scores well on all four. A chat window on a consumer account, used by whoever is at the desk, scores poorly on most of them, which is why the same cleaned paragraph can pass in one and fail in the other.
For a disability provider, our NDIS guide takes both steps through an invented progress note.
Two Australian datasets that were called de-identified
In Australia the line between de-identified data and pseudonymised data shows up most clearly in the Commissioner’s own investigations rather than in the guidance, and one of them reached opposite conclusions about different people in the same release.
MBS and PBS, 2016: the provider numbers came back
On 1 August 2016 the Department of Health published on data.gov.au claims data for a 10% sample of people who had claimed Medicare benefits since 1984 or pharmaceutical benefits since 2003. Each patient appeared under an identification number, years of birth stood in for dates of birth, and service dates were shifted by up to 14 days either way.
Medicare provider numbers were handled differently. They were encrypted with a method that, the OAIC later found, was reversible. On 8 September 2016, researchers Chris Culnane, Benjamin Rubinstein and Vanessa Teague of the University of Melbourne told the Australian Bureau of Statistics they had decrypted them. The dataset came down that day, 38 days after it went up.
One dataset, two answers
The Commissioner’s findings split along exactly that line. For patients, re-identification would have required unusual features or a very full knowledge of someone’s medical history, and the Commissioner concluded that patients in the dataset were not reasonably identifiable. For them, the release was de-identified.
For providers, a unique identifier “can be derived from the dataset following a process of decryption”. That made their information personal information, and the Commissioner found the Department breached APP 6 by disclosing it, and APPs 1 and 11 in preparing the dataset. An enforceable undertaking was accepted on 23 March 2018. The purpose test behind that APP 6 finding, applied to AI tools, gets its own treatment in the guide to AI under the Privacy Act.
| in the 2016 release | how they were coded | could the code be reversed? | what the Commissioner found |
|---|---|---|---|
| Patients | an identification number, year of birth, service dates shifted by up to 14 days | only with unusual features or a very full knowledge of someone’s medical history | not reasonably identifiable: de-identified |
| Medicare providers | provider numbers encrypted | yes, the encryption was reversible | personal information: APP 6 breached |
A reversible code is pseudonymisation. When the key, or the method that works as one, is within reach of the recipient, the code is just the identifier in a different font.
Flight Centre, 2017: checked, and still disclosed
The second case is smaller and closer to a working office. In 2017 Flight Centre gave 16 teams at a design jam access to a customer dataset, despite preliminary checks to de-identify or remove personal information. The data included credit card and passport details, and the error was found after it had been available for 36 hours.
The Commissioner determined in December 2020 that Flight Centre interfered with the privacy of almost 7,000 customers, breaching three APPs. The lesson for anyone about to paste a file is not about hackathons. Checking a file and believing it clean are the same experience right up until someone else reads it. When that someone is a chatbot, the first hours after the mistake are the subject of a separate breach guide.
I-MED: what it took for the OAIC to accept the data as de-identified
The case that went the other way is the one most relevant to AI, because it was about training a model. On 19 September 2024 media reports alleged that I-MED Radiology Network had disclosed patient scans to Annalise.ai, a former joint venture with Harrison.ai, to train a diagnostic AI. The OAIC opened preliminary inquiries the next day.
In its report published on 31 July 2025, the Commissioner was satisfied that the patient data had been “de-identified sufficiently that it was no longer personal information for the purposes of the Privacy Act”, and closed the inquiries without further action.
What I-MED actually did
The scale is what makes the method worth reading. Between April 2020 and January 2022, I-MED shared under 30 million patient studies and a similar volume of diagnostic reports. The OAIC’s account of what happened before any of it left is specific, and it reads nothing like a label swap.
| what I-MED had in place | what a paste into a chatbot usually has |
|---|---|
| Two hashing techniques for IDs, names, addresses and phones | a label on the name |
| Dates shifted to a random date within set years | the real dates |
| Outliers aggregated into large cohorts | the rare detail left in |
| Text near the edge of each image redacted | the free text untouched |
| A contract forbidding re-identification | whatever the account’s terms say |
| Separate environments and secure storage | the provider’s own systems |
Even with all of that, Annalise.ai reported a very small number of occasions where personal information had reached it in error through failures in the process, and that material was deleted or de-identified. The Commissioner still accepted the outcome, because the risk had been reduced to a sufficiently low level and was backed by governance.
The part a busy practice cannot copy
The right column is no criticism of anyone; that is simply what a prompt looks like. A practice cannot shift the dates in a client’s matter without losing the point of asking, and it cannot bind a consumer chatbot to a contract of its own drafting. What those accounts do promise is set out for ChatGPT and for Copilot.
The OAIC was also careful to add that the case study “should not be taken as an endorsement of I-MED’s acts or practices”. It shows that de-identification for AI is achievable in Australia, not that a label on a name achieves it.
Is pseudonymised data enough to put client data into ChatGPT?
For most client work, no. The problem lies in everything the label does not touch: the town, the business, the family member, the dates, the amounts and the story that ties them together. If those still point to one person, the prompt carries personal information and the paste is a disclosure. In litigation the NSW Supreme Court asks a different question, and its practice note judges subpoenaed material by the platform, not by the labels.
The OAIC’s October 2024 guidance on AI products bought off the shelf spells out the consequence with a worked example. Staff at an insurer paste a customer’s claim, health details and all, into a public chatbot to draft an assessment. By doing so, the OAIC says, the insurance company “is disclosing the information to the owners of the chatbot”, and APP 6 applies to that disclosure.
What pseudonymising does buy you
It is not wasted effort, and the regulator says so. The same guidance says the OAIC “expects organisations to consider what information is necessary and assess whether there are ways to minimise the amount of personal information that is input into the AI product”, and suggests techniques that protect privacy while doing it.
Pulling the TFN, the Medicare card number, the contact details and the client’s name out of a prompt is exactly that minimisation. It means a leak would carry less, it removes the identifiers that carry their own legal regimes, and it makes the remaining judgement smaller. That is a genuine gain, and it is the claim worth making. For a clinic the Medicare number is the clearest case, because a principle of its own governs it.
What it does not change
It does not change the legal status of a file that still identifies someone. APP 6 still asks whether the disclosure fits the purpose you collected the information for. APP 8 still applies if the recipient is overseas. And the regulator’s advice on public chatbots is unchanged: it “recommends that organisations do not enter personal information, and particularly sensitive information, into publicly available generative AI tools”.
Whether one particular paste also triggers the notifiable data breach scheme is another matter, with a deadline of its own, worked through in our guide on putting client data into a chatbot. Much depends on the type of account, and our guides to Claude, Gemini and Copilot set out what we found for each.
A client note before and after, on version 0.2.8
Principles are easier to judge with a real example in front of you. The note below is the kind of thing a bookkeeper or a tax agent might want help drafting a reply to. Everything in it is invented, the numbers included: the tax file, Medicare and ABN numbers were generated to pass their checksums.
BEFORE Client: Marguerite Halloran, TFN 609 905 481, DOB 02/11/1968.
Medicare 6689 19577 9. Mobile 0491 570 156, m.halloran@example.com.
14 Wattlebird Cres, Myrtleford VIC 3737. ABN 44 148 175 704.
She sold the only vet clinic in Myrtleford in March and wants the
capital gain split with her brother Tobias before the June lodgement.
AFTER Client: [PERSON_1], TFN [TFN_1], DOB [BIRTH_DATE_1].
Medicare [MEDICARE_1]. Mobile [PHONE_1], [EMAIL_1].
[ADDRESS_1], [ADDRESS_2]. ABN [ABN_1].
She sold the only vet clinic in Myrtleford in March and wants the
capital gain split with her brother Tobias before the June lodgement.
The AFTER block is the output Nonimo’s engine produced from that note on 23 September 2026, in the 0.2.8 release for both Mac and Windows, with its suggestions accepted and nothing edited by hand.
Look at the last two lines, which are identical before and after. The only vet clinic in one named town, sold in one named month, by a woman with a brother called Tobias. Anyone in Myrtleford could name her. The file is pseudonymised. It is nowhere near de-identified, and no tool could have made it so without deleting the question.
The sentence that decides it is the one you would have to rewrite
The fix is a human edit, and it is usually quick: “a client sold a regional veterinary practice this financial year and wants a capital gain split with a sibling”. The accounting question survives intact. The client does not.
That rewrite is the second step in the OAIC’s process, done by the one person who knows which details matter. The sector version of the same exercise, for progress notes, is in our NDIS guide, and the pages for accountants and law firms show the identifiers each profession handles most.
De-identified vs pseudonymised in your AI policy
The difference between the two words matters most in the one document an auditor, an insurer or a client will actually read. A policy that promises de-identification commits you to a legal result. A policy that describes pseudonymisation commits you to a method, which you can show you followed. The same care applies to the data paragraph of a client engagement letter, where the wrong word becomes a promise to a client.
Say that identifiers are replaced before anything is entered, and that the key stays on the practice's own systems. Do not call the result de-identified.
Somebody has to check the remaining text for the details that identify, and that person should be named, not implied.
A key uploaded alongside the text turns a pseudonymised prompt back into an identified one.
A personal login and a company account give different answers.
The tool, the date, the approver and the reasons.
The security side of this has a date attached. APP 11.3, which applies to information held after 11 December 2024, says the reasonable steps an entity takes to secure personal information include technical and organisational measures. Swapping identifiers out before a prompt is sent counts as a technical measure. The person who reads the story is an organisational one.
If a file is only held for a job that is finished, the Act adds another option you may have forgotten. Under APP 11.2, personal information you have no further permitted use for must be destroyed or de-identified, unless a law requires it to be kept. A file that no longer exists cannot leak. It is also one less thing to weigh when you read what cyber insurance covers.
Where Nonimo fits, and which of the two words it can claim
Nonimo works locally, on the Mac or Windows machine where the file already sits. Before a prompt is sent anywhere, it swaps the identifiers it recognises for labels, lists every swap so a person can reverse any of them, and stores the table linking each label to its value, encrypted, on that same machine.
For Australian practices the recognised set includes the TFN, ABN, ACN and Medicare number, and the Australian home page shows how it works on a practice computer.
That makes it a pseudonymisation tool, and nothing grander. In Privacy Act terms, the note it hands back is still personal information for as long as you hold the table, and deciding what the story gives away remains with whoever knows the client.
What it keeps on the machine is set out on Nonimo’s security page, and the licence terms on the licence page. The guides on using AI with client data are written so that you can weigh tools like ours as well.
Before the paste: one question, asked about the reader
Forget whether the labels went on. The useful question under the Act is about the next reader: holding only the text you plan to send, could they work out who your client is?
A yes, or a shrug, means the text is still personal information and sending it is a disclosure, labels or no labels. The labels were not wasted. They stripped out the numbers that bring their own legal trouble, and shrank what is left to a decision one person can make well. The decision after that, about the tool and the account, begins in the Privacy Act guide.
Sources
Every page below was open in front of us on 23 September 2026.
- Privacy Act 1988 (Cth), Federal Register of Legislation, compilation of 4 June 2026, in which the words pseudonymise and pseudonymisation do not appear. The section 6(1) definitions of personal information and de-identified, and Australian Privacy Principles 2, 6, 8 and 11, including APP 11.2 on destroying or de-identifying information no longer needed and APP 11.3 on technical and organisational measures.
- OAIC, De-identification and the Privacy Act, publication date 21 March 2018, with a notice that it is being updated for the Privacy and Other Legislation Amendment Act 2024. The two steps of de-identification, the warning that removing direct identifiers alone may not be enough, the description of encryption and hashing as obscuring rather than removing, the custodian example, the test in the hands of the other entity, the four parts of a data access environment, the motivated intruder minimum, and the note that APP 11.3 applies to information held after 11 December 2024.
- OAIC, APP Guidelines chapter 2, APP 2 Anonymity and pseudonymity, version 1.1, 22 July 2019. The description of a pseudonym and the statement that its use does not necessarily mean an individual cannot be identified.
- OAIC, Guidance on privacy and the use of commercially available AI products, published 21 October 2024. The insurer example in which entering a claim into a public chatbot is a disclosure to its owners, the expectation that personal information input into an AI product be minimised, and the best practice recommendation not to enter personal information into publicly available generative AI tools.
- OAIC, MBS/PBS data publication, Commissioner initiated investigation report. The publication on 1 August 2016, the 10% sample, the date shifting of up to 14 days, the reversible encryption of provider numbers, the report by the University of Melbourne researchers on 8 September 2016, the finding that patients were not reasonably identifiable while providers were, the breaches of APPs 1, 6 and 11, and the enforceable undertaking of 23 March 2018.
- OAIC, Report into preliminary inquiries of I-MED, 31 July 2025. The media reports of 19 September 2024, the inquiries opened on 20 September 2024, the under 30 million studies shared between April 2020 and January 2022, the six processing steps and the contractual terms, the small number of errors reported by Annalise.ai, the Commissioner’s conclusion, and the statement that the case study is not an endorsement.
- OAIC, Flight Centre found to have interfered with privacy, 7 December 2020. The 2017 design jam with 16 teams, the preliminary checks to de-identify, the credit card and passport details, the 36 hours, the almost 7,000 customers and the three APPs breached.
The client note in this guide is invented, including the town’s only vet clinic. The name, street, email, phone and clinic belong to nobody we know of, and the tax file, Medicare and ABN numbers were generated to pass their checksums.
Common questions
What is the difference between de-identified and pseudonymised data?
De-identified data has a legal status, and pseudonymised data has simply been through a technique. Under section 6(1) of the Privacy Act, information is de-identified when the person is no longer identifiable or reasonably identifiable, and the Act stops applying to it. Pseudonymised means the identifiers were swapped for labels and a way back was kept. The Act does not use that word, and whoever holds the way back is still holding personal information.
Does the Privacy Act define pseudonymisation?
No. The Act defines personal information and de-identified, and nothing in between. The nearest word, pseudonymity, belongs to APP 2, and it means something else: the right of a customer to deal with you under a name that is not theirs. The OAIC's APP 2 guidance adds that using a pseudonym does not necessarily mean the person cannot be identified.
Is pseudonymised data still personal information in Australia?
For the business that holds the key, generally yes. The OAIC's de-identification guidance says a dataset may be personal information when handled by the custodian who kept the original, and de-identified when handled by a different entity. So the same labelled file can be personal information on your desk and, if nothing else in it identifies anyone, something less at the other end.
Can I put pseudonymised client data into ChatGPT?
Only if nobody is reasonably identifiable from what arrives, and a client file rarely clears that bar once the story is left in. If someone still is, the paste is a disclosure under APP 6, and possibly an overseas one under APP 8. The OAIC's best practice advice is to keep personal information out of public generative AI tools altogether, and it expects whatever does go in to be kept to a minimum.
What did the OAIC decide about I-MED and de-identified data?
In a report published on 31 July 2025, the Commissioner was satisfied that patient data I-MED shared with Annalise.ai to train a diagnostic model had been de-identified sufficiently to fall outside the Privacy Act. I-MED had hashed identifiers, shifted dates, aggregated outliers, redacted text near images and bound the recipient by contract. The OAIC said the report is not an endorsement of I-MED's wider practices.
Is encrypting or hashing a client number enough to de-identify it?
No, not by itself. The OAIC lists encryption and hashing as techniques that obscure an identifier rather than remove it. In 2016 the Department of Health released Medicare data with provider numbers encrypted, researchers at the University of Melbourne reversed the encryption, and the Commissioner found the release breached APP 6 for those providers.
What does reasonably identifiable mean under the Privacy Act?
It asks whether identification is both technically possible and reasonably likely, given the information itself, what else the person holding it can reach, and how practical the linking would be. The OAIC sets a minimum test: whether a reasonably competent, motivated person with no specialist skills, internet access and public documents could identify someone. Context decides, so the answer can change with the reader.
Do I have to de-identify client information I no longer need?
Destroy or de-identify, one or the other. APP 11.2 says that when an APP entity no longer needs personal information for any permitted purpose, it must take reasonable steps to destroy it or de-identify it, unless a law requires it to be kept. For a practice, that usually means the retention periods your professional rules and the tax law already set.
Does Nonimo de-identify documents?
No, it pseudonymises them. It replaces identifiers with labels on your computer before the text goes anywhere, shows you what it changed, and keeps the way back encrypted on that same computer. Under the Privacy Act the file stays personal information in your hands. Whether what you send is identifiable at the other end is a judgement about context that remains yours.