Pseudonymisation under UK GDPR: what to take out of a client file
· Updated · Written and maintained by Joaquín Trapero, Nonimo
You have a client file open and a chat window next to it. Before the file goes across, something has to come out of it, and the question that matters is less how much you take out than what the regulator will call the thing you are left holding.
It will almost certainly be pseudonymised, not anonymous, and where pseudonymisation ends and anonymisation begins decides whether the UK GDPR still applies. The ICO has written the reason down twice, in two different places, in almost the same words: removing the direct identifiers is not enough.
This guide is about applying that test to one document, today, under UK law. If what you want is the vocabulary itself, and the way the three words differ across regimes, that lives in our guide on the difference between deidentified and anonymised data. This one assumes you already have a file and a deadline.
Pseudonymisation under the UK GDPR: what the word buys you
Article 4(1)(5) of the UK GDPR defines pseudonymisation as processing personal data “in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately”. The ICO puts the same idea in working language: techniques that “replace, remove or transform information that identifies a person and store it separately”.
Read the second half of that definition rather than the first. The link is filed elsewhere rather than destroyed, and somebody still holds the drawer it is filed in. That somebody is usually you.
What it buys, and what it does not
So you get a genuine reduction in risk and a recognised security measure. What you do not get is an exit. Anonymisation is the exit, and it is a different standard entirely.
| what you did | what changes for you | |
|---|---|---|
| Masking | hid part of a value you kept | nothing legally. It is a display technique |
| Pseudonymisation | replaced the identifier, kept the key | risk falls, obligations stay |
| Anonymisation | made identification no longer reasonably likely | the UK GDPR stops applying |
The practical consequence is that a cleaned file is still a file you answer for. Every question in our guide to what a UK business has to check in a GDPR compliant AI tool continues to apply to it, unchanged.
Why the ICO will not call your cleaned file anonymous
The regulator has been unusually direct about this, and it matters because the sentence is short enough to quote back to anyone in your office who thinks the job is done.
The sentence the regulator has now written twice
In its anonymisation guidance, the ICO says:
“Simply removing direct identifiers from a dataset is insufficient to ensure effective anonymisation. If it is possible to link someone to information in the dataset that relates to them, then the data is personal data.”
It then says it again, in a different place and for a different audience. In its innovation advice, last updated on 16 December 2025, the answer to a question about employee data reads: “The removal of direct identifiers such as a name or an identification number is insufficient to ensure effective anonymisation.”
Two separate publications, one position. There is no reading of the ICO’s guidance in which stripping the name and the reference number produces an anonymous document.
Anonymisation is a destination, and you rarely reach it
The guidance was published on 28 March 2025, and it carries a notice that it is under review because of the Data (Use and Access) Act 2025.
That is worth knowing before you build a process on any single paragraph of it. It is the same review that has left the ICO’s older AI guidance flagged as under review too, which we go through in what a UK business has to check.
What the guidance also does is refuse the binary.
“In practice, identifiability may be viewed as a spectrum that includes the binary outcomes at either end, with a blurred band in between.”
Most real client files sit in that blurred band, which is precisely why the label you give them matters.
One more thing that surprises people, and it is in the same guidance. Anonymising is itself regulated: “applying anonymisation techniques to turn personal data into anonymous information counts as processing personal data”. The cleaning step needs a basis of its own. It is not a preparatory act outside the regime.
Is pseudonymised data personal data? The ICO answers in one sentence
Its pseudonymisation page asks the question as a heading and answers it immediately.
“Pseudonymised data is personal data in the hands of someone who holds the additional information.”
Everything turns on the qualifier at the end. Identifiability here is relative to who is holding the file. The same document can be personal data for your firm, because you keep the mapping, and much closer to anonymous for a reader who has no way to reach it.
That is why the ICO’s practical requirements are about custody rather than about the substitution itself. It asks organisations to hold the additional information separately and keep it secure, and to “store the additional information and pseudonymised data in distinct physical locations (eg using separate databases or network segmentation)”. A key kept in the same folder as the file is not a key.
The European rulings people cite, and the weight they carry here
Search this question and the first page fills with commentary on recent Court of Justice case law about when pseudonymised data is personal data. Those rulings are worth reading. In the UK they are persuasive rather than binding, and that matters.
Section 6 of the European Union (Withdrawal) Act 2018 is the provision to know. A court or tribunal, it says at subsection (1)(a):
“is not bound by any principles laid down, or any decisions made, on or after IP completion day by the European Court”
Subsection (2) then allows a court to “have regard to anything done on or after IP completion day by the European Court”. So a Luxembourg judgment handed down after Brexit is something a UK court may weigh, not something it must follow.
If you are deciding what to do with a file this afternoon, the ICO’s published position is the one that will be applied to you. The wider question of whose law reaches your use at all is a separate one, and we cover it in our guide on the EU AI Act outside the EU.
The motivated intruder, and what the ICO allows them
This is the test that decides, and it is the most useful thing the ICO has given UK firms, because it converts a legal abstraction into a question a person can actually answer at a desk.
A motivated intruder, in the ICO’s words, is this.
“Someone who wishes to identify a person from the anonymous information that is derived from their personal information.”
You are asked to assume they are “reasonably competent” and that they have “access to appropriate resources (eg the internet, libraries, public documents)”. You are also asked to assume they will make enquiries, including “advertising for anyone with that knowledge to come forward”.
What you are not asked to assume is a state actor with unlimited time. The standard is what is reasonably likely, not what is theoretically conceivable, and the ICO says a “purely hypothetical or theoretical chance of identifiability” does not need to be considered.
Singling out, linkability and inference
The guidance splits the risk into three, and they fail in different ways. Running a file past all three takes about a minute.
| the ICO’s description | what it looks like in a file | |
|---|---|---|
| Singling out | isolating the records that relate to one person | one matter, one date, one town |
| Linkability | combining records, “the mosaic or jigsaw effect” | your file plus a public register |
| Inference | guessing details about someone already identifiable | the diagnosis implied by the treatment |
The order matters less than the fact that they are three, not one. A file can pass as unidentifiable on the name and fail on the shape: it is the only matter of its kind on your books, and a public register closes the last gap.
Councils meet the same three tests from the other direction when they answer a freedom of information request, which is why they appear again in our guide for council staff.
What British office paper actually contains, measured
Before deciding what to strip, it helps to know what is actually printed on the documents in front of you. We measured it, on public British material rather than on assumptions.
The corpus was 23 public British documents totalling 92,217 words: six blank official forms from gov.uk, nine judgments from Find Case Law, three ICO guides, three NHS pages and two pages of the HMRC manual. Every document has a recorded provenance and a checksum.
The label is everywhere and the value is almost nowhere
The result was the opposite of what we expected, and it changes how a UK file should be read.
| the label appears | a real value appears | |
|---|---|---|
| NHS number | 57 times, in 3 documents | 0 |
| National Insurance | 41 times, in 7 documents | 0 |
| Postcode | 31 times, in 7 documents | 4 |
| Driving licence | 1 time, in 1 document | 0 |
No vehicle registration and no mobile number of the form British numbers take appeared anywhere in the corpus either. British public paper carries the field name and almost never the entry, because the forms are blank and the judgments are already anonymised.
One caveat belongs with that. 46 of the 57 appearances of the NHS number label sit on a single page, and one document accounts for 34.1 per cent of the whole corpus, so the label counts are more concentrated than the totals suggest.
What that means for the file on your desk
Your file is the other kind of document. It is the completed form, the attendance note, the letter of instruction, and it carries the values that the public corpus does not. The searching you do for names is the easy half. What tends to survive is everything that was never an identifier in the first place: the sequence of dates, the amount, the employer, the court, the fact that there is only one matter like it.
That is also why the account you use matters as much as the cleaning. Our guides on what ChatGPT does with your data in the UK and on whether Copilot is safe for confidential information go through the settings that decide where a cleaned file ends up.
Run the test on one paragraph
Here is an attendance note of the kind that goes into a chat window twenty times a week, in whichever assistant your firm has settled on, from Gemini to Claude. The names and the numbers are invented.
Attending Mrs J Tredannick at the Marlbrook office, 14 March. Client is a theatre sister at the district general hospital, off sick since the fall on 2 January. NHS number 000 000 0000. Instructing us on a claim against her employer. Previous solicitors Weatherby and Crane, our reference WC/8840.
The same paragraph with only the number gone
Take out the NHS number and the surname, which is what most people do, and read what remains.
Attending our client at the Marlbrook office, 14 March. Client is a theatre sister at the district general hospital, off sick since the fall on 2 January. Instructing us on a claim against her employer. Previous solicitors Weatherby and Crane, our reference WC/8840.
What is still doing the identifying
Nothing in the second version is a direct identifier, and it singles a person out comfortably. One hospital in the town, a specific role inside it, a dated absence and a claim against the employer will narrow this to one individual for anyone who works there. The previous solicitors’ reference is a second route: it is a key held by a third party, and the motivated intruder is allowed to make enquiries.
| what is left | the ICO test it fails | why it still identifies |
|---|---|---|
| theatre sister at the district general hospital | singling out | one hospital in the town, one role inside it |
| off sick since the fall on 2 January | singling out | a dated absence |
| a claim against her employer | singling out | obvious to anyone who works there |
| Weatherby and Crane, reference WC/8840 | linkability | a key held by a third party, who can be asked |
Apply the ICO’s three tests and the file fails on singling out and on linkability, with every name removed. That is no failure of effort, just what the regulator means by the blurred band, and it is the reason the label on your file is pseudonymised. The fall behind her sick leave is health data in its own right, since the ICO treats even a broken leg as health data.
What the ICO told organisations to do before sharing a document
There is a direct answer to the question people actually ask, and it is short. In the same innovation advice, the ICO says this.
“If an organisation is able to anonymise the information, or remove identifiable information from the documents shared, then they should do so.”
Read what it does and does not say. It is an instruction to reduce what you send. It is not a statement that reducing it makes the processing lawful, and it does not say the result will be anonymous. The lawful basis, the transparency and the contract all remain to be dealt with, which is the ground our AI policy template is built to cover.
The practical reading for a firm is that doing nothing is not neutral. The regulator has said in writing that you should remove what you can. Sending the whole file because the cleaned version would still be personal data anyway is an argument that does not survive contact with that sentence.
For a solicitor, the SRA adds a second test
Identifiability is a data protection question. Confidentiality is a professional one, and in England and Wales it is owed to the client whether or not the material is personal data at all. It reaches the 179,054 solicitors who held a practising certificate at the end of August 2026. Since it is owed to the client, the client’s written agreement can cover a tool, through an AI clause in your terms of business.
On 17 August 2026 the SRA published a warning notice, Misuse of AI, tying AI use to paragraph 6.3 of both Codes of Conduct. Two sentences in it close arguments that get made in real offices. The first: “Both free to use and paid for AI systems may pose risks to client confidentiality.” Paying does not settle it.
The safeguards sentence, and the question it forces
The second is the operative one: “Client information should only be entered into AI systems where appropriate contractual, technical and organisational safeguards are in place to protect confidentiality.”
Note the order. Contractual comes first, and a cleaning tool is not a contract. Pseudonymising a file is a technical safeguard, and on the SRA’s wording it is one of three things you need, not a substitute for the other two. Applied on the fee earner’s own machine before a witness statement reaches ChatGPT or Claude, it is the safeguard a solicitor’s practice can put on the desk itself.
Litigators meet the same duty in the Technology and Construction Court’s own guide, where keeping the underlying data confidential sits beside what you tell the judge about AI.
Of the 8,923 firms the SRA regulates, 1,321 are sole practitioners, which is 14.8 per cent of them. Those proportions are the reason this guide is written the way it is.
In those 1,321 firms the person deciding whether the file is clean enough is also the person on the file, with no second reader and no data protection officer of their own. The same is true of most of the councils and small teams that the local government guide is written for.
Section 171, and why the ICO avoids the word deidentified
British writing about this borrows “deidentified” from American practice, and it causes a specific problem here, because the word does appear in UK law, once, and it means something narrow.
Section 171 of the Data Protection Act 2018 makes it “an offence for a person knowingly or recklessly to re-identify information that is de-identified personal data without the consent of the controller responsible for de-identifying the personal data”.
The Act defines de-identified as processed so that it “can no longer be attributed, without more, to a specific data subject”. It then provides defences, including where undoing it “was necessary for the purposes of preventing, investigating or detecting crime”, where it was required or authorised by law, and where “in the particular circumstances, was justified as being in the public interest”.
A state with an offence attached, not a status that releases you
The trouble starts in your own paperwork. Using the word as a synonym for anonymous, in a file note or in a client letter, describes your own document as something the law does not agree it is. The comparison across the three regimes is set out in our guide to the vocabulary.
Where cleaning the file stops being the answer
Three things sit outside what any amount of removal can fix, and they are worth naming because tools in this category are often sold as though they fix all three.
- Confidentiality is not identifiability. A document with every identifier gone can still disclose a client’s affairs. The SRA duty and the UK GDPR are two separate tests, and a clean pass on one is not a pass on the other.
- Sending is still a disclosure. Whether the transfer to a provider is lawful depends on the contract, the account type and the transfer route, not on how much of the text you removed first.
- Evidence is a separate matter from protection. Knowing that you cleaned a file is not the same as being able to show, months later, what was in it and what came out.
If the worst has already happened, the question changes from prevention to notification, and that is covered in our guide on whether putting client data into ChatGPT is a breach. The insurance angle, which is a different conversation again, is in our comparison of cyber cover for UK small businesses.
What ours does in the UK, and the loose end we publish
Nonimo replaces identifiers in a document before you send it. It runs on your machine, the mapping back to the person is kept encrypted there, and the substitution is therefore reversible by design. By the definitions at the top of this guide, that is pseudonymisation, and it leaves your file inside the UK GDPR.
It is worth being exact about the British layer, because it behaves differently from the others. Nothing it finds is substituted in silence.
Nothing is substituted silently here, and why
Where a value has an unmistakable shape (an IBAN, a card number, an email address, a travel document), the substitution happens without asking. The British national identifiers do not qualify, and the arithmetic is the reason.
An NHS number is ten digits with no fixed prefix, and its check digit is a modulus 11 calculation, so roughly 1 in 11 arbitrary strings of ten digits will satisfy it. That is a useful filter and it is nowhere near an identification. So every British national identifier the tool finds is shown to you, with the reason it was flagged, and every change can be undone. You see the change and you decide.
The loose end, and we would rather you heard it from us
A number with no label beside it can be missed. The detection for NHS numbers and UTRs requires the field name to be present, which is a deliberate choice made to avoid false positives, and it has a cost: an NHS number sitting alone in a spreadsheet column, with no heading, is left as it is.
The reason is the corpus above: it held no real values to measure a rule that works on shape alone, so the detection asks for the field name instead.
What it does not do, said plainly
By default nothing is blocked. It does not read scanned images. It does not make you compliant with the UK GDPR, which is a property of an organisation rather than of a product, and no supplier can sell it to you. There is a free plan with 200,000 words a month. Beyond the single desk, a practice rolls out the managed extension centrally and watches the administrator dashboard.
Why we do not publish a single accuracy figure
A single percentage is unfalsifiable without the thing that sits beside it. A tool that flags everything has perfect recall and is useless, and one that flags almost nothing has no false positives and is worse. We publish both numbers together or we publish neither, and for the British layer the corpus above is why neither exists yet. If you want the case for asking a supplier for evidence rather than taking them at their word, our page for organisations is where we make it.
The question to ask before you paste
Not “have I taken enough out”. The ICO’s question is better, and you can answer it in the time it takes to read the paragraph again.
Could a reasonably competent person, with the internet and a willingness to ask around, work out who this is from what is left? If the answer is yes, or if the most honest answer you can give is probably not, the document is pseudonymised. The UK GDPR still applies to it, and every obligation you had before the cleaning you still have after it.
That is not an argument for sending the file uncleaned. Removing what you can is what the regulator has told you to do. It is an argument for being accurate about what you achieved, because the label you write in the file note is the one you will be held to.
Two documents make that label stick.
- A written position on which tools your firm allows, and on what may go into them. That is what our AI policy template is for.
- An honest answer to your insurer, before the renewal rather than after the incident. The questions themselves are in our guide to the AI questions on an insurer’s proposal form.
Sources
Checked 21 September 2026.
- ICO, Pseudonymisation. The definition in working language, the answer that pseudonymised data is personal data in the hands of someone who holds the additional information, and the requirement to store the additional information and the pseudonymised data in distinct physical locations.
- ICO, How do we ensure anonymisation is effective?. The motivated intruder and what is assumed about them, the statement that simply removing direct identifiers is insufficient, the identifiability spectrum, and the definitions of singling out, linkability and inference.
- ICO, About this guidance. Published 28 March 2025, and the notice that the guidance is under review because of the Data (Use and Access) Act.
- ICO, Introduction to anonymisation. That applying anonymisation techniques counts as processing personal data, and the distinction between anonymisation and pseudonymisation.
- ICO, Innovation advice: previously asked questions, last updated 16 December 2025. That organisations able to anonymise or remove identifiable information from documents shared should do so, and that removing direct identifiers such as a name or an identification number is insufficient.
- ICO, Data (Use and Access) Act 2025. Royal Assent on 19 June 2025, and the confirmation that the provisions affecting data protection law are in force.
- UK GDPR, Article 4, legislation.gov.uk. The Article 4(1)(5) definition of pseudonymisation quoted above.
- Data Protection Act 2018, section 171, legislation.gov.uk. The re-identification offence, the definition of de-identified personal data, and the defences.
- European Union (Withdrawal) Act 2018, section 6, legislation.gov.uk. Subsection (1)(a), that a court or tribunal is not bound by principles laid down or decisions made by the European Court on or after IP completion day, and subsection (2), that it may have regard to them.
- SRA, warning notice Misuse of AI, 17 August 2026. Both quoted sentences, and paragraph 6.3 of the Codes of Conduct.
- SRA, Population of solicitors. 179,054 practising solicitors at the end of August 2026.
- SRA, Breakdown of solicitor firms. 8,923 regulated firms in August 2026, of which 1,321 sole practitioners, 14.8 per cent of the total.
- Nonimo corpus of British public documents, measured 25 August 2026. 23 documents and 92,217 words, with a manifest and a checksum for each: six blank official forms from gov.uk, nine judgments from Find Case Law, three ICO guides, three NHS pages and two pages of the HMRC manual. Source of the label counts, of the four postcodes, of the concentration caveat and of the loose end on unlabelled values.
Nonimo is the software that does this on your own computer: it masks client names and IDs before your text reaches ChatGPT . No account, and your client's details never leave your machine.
Common questions
What is pseudonymisation under the UK GDPR?
Processing that means personal data can no longer be attributed to a specific person without additional information, where that additional information is kept separately and secured. That is Article 4(1)(5) of the UK GDPR, and the ICO describes it as replacing, removing or transforming identifying information and storing it separately.
Is pseudonymised data still personal data?
Yes, in the hands of anyone who holds the additional information. The ICO answers the question in those words on its pseudonymisation page, because data protection law treats information as personal data whenever a person is identified or identifiable, directly or indirectly.
Does removing names and reference numbers anonymise a document?
Not on its own. The ICO says simply removing direct identifiers is insufficient to ensure effective anonymisation, and that if someone can still be linked to information that relates to them, the data is personal data. The file stays inside the UK GDPR.
What is the motivated intruder test?
The ICO's way of asking whether identification is reasonably likely. You assume someone reasonably competent who wants to identify a person, with access to the internet, libraries and public documents, and who will make enquiries. You then ask whether they would succeed.
Does the UK GDPR use the term de-identified?
Only in one place. Section 171 of the Data Protection Act 2018 makes it an offence to knowingly or recklessly re-identify de-identified personal data without the controller's consent. The ICO does not treat de-identified as a synonym for anonymous.
What does the ICO say about putting client documents into generative AI?
In its innovation advice, last updated on 16 December 2025, the ICO says that if an organisation is able to anonymise the information, or remove identifiable information from the documents shared, then they should do so. It does not say that doing so makes the use lawful.
Do solicitors have an extra duty beyond the UK GDPR?
Yes. The SRA warning notice Misuse of AI, published on 17 August 2026, ties AI use to paragraph 6.3 of both Codes of Conduct and says both free to use and paid for AI systems may pose risks to client confidentiality. Confidentiality is owed whether or not the file is personal data.
Is anonymising a document itself regulated?
Yes. The ICO states that applying anonymisation techniques to turn personal data into anonymous information counts as processing personal data. You need a lawful basis for the cleaning step, not only for whatever you do with the cleaned file afterwards.
Can a client file still identify someone after every name is gone?
Often, yes. The ICO describes singling out, linkability and inference as three separate routes, and calls linkability the mosaic or jigsaw effect. A role, a date, a town and an amount can narrow a file to one person without a single name in it.