assurance · Level 3
Privacy and data minimisation in AI applications
UK data protection basics for anyone building or buying an AI feature, with a small redaction pipeline you run and test yourself.

Start with the essentials
The short answer
Data minimisation means an AI feature uses only the personal data it needs for a stated purpose and keeps it no longer than necessary. Personal data hides in prompts, logs, embeddings and outputs as well as databases. The ICO says most uses of AI need a data protection impact assessment. Redacting prompts reduces risk but does not remove it, and redacted data usually remains personal data.
What you will learn
- You will be able to find personal data in every part of an AI feature, including prompts, logs, embeddings and outputs.
- You will be able to apply the UK GDPR principles, lawful basis, transparency and rights to an AI feature in plain words.
- You will be able to screen an AI project for a DPIA, name the controller and processor, and spot a restricted international transfer.
- You will have built and tested a redaction and minimisation pipeline, and be able to explain what it misses and why.
- You will be able to write a record of processing entry and a retention rule, and review a data flow for privacy problems.
Who it is for
People building or buying an AI feature in a UK organisation: developers, product owners, analysts, and IT or procurement staff who must keep personal data under control. You need to read short Python and run a script from a terminal.
Before you start
- A plain-English idea of how chatbots work, as in What is AI?, and enough Python to run a script from a terminal, as in Python: from first script to a useful automation. Stay safe with AI covers your rights as an individual; this workbook takes the organisation's side.
Read a sample · Chapter 04 of 06
Build and test a redaction pipeline
Drop the fields you do not need, then redact identifiers from the free text that remains.
Article 25 requires data protection by default: only personal data necessary for each specific purpose should be processed. Allowlisting keeps named fields and drops the rest, even fields added later. Redaction swaps identifiers in free text for labels such as [EMAIL], using regular expressions (text patterns) and a list of known names. It works on language, so it guesses.
Save this as redact.py in a new folder and run it (use python3 on macOS and Linux). It needs only the standard library and was tested with Python 3.12.10 on Windows 11.
"""redact.py: minimise and redact a ticket before it goes to an AI service. All data invented."""
import re
KEEP = ["ticket_id", "category", "message"] # the field allowlist
NAMES = ["Ada Fictional", "Sam Sampleton", "Jo Testperson"] # invented customers
PATTERNS = {
"EMAIL": r"[\w.+-]+@[\w-]+(?:\.[\w-]+)+",
"PHONE": r"(?:\+44\s?|\b0)\d{2,4}\s?\d{3,4}\s?\d{3,4}\b",
"NINO": r"\b[A-Z]{2}\s?\d{2}\s?\d{2}\s?\d{2}\s?[A-D]\b",
"POSTCODE": r"\b[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}\b",
}
def minimise(ticket):
return {k: v for k, v in ticket.items() if k in KEEP}
def redact(text):
counts = {}
rules = [("NAME", re.escape(n)) for n in NAMES] + list(PATTERNS.items())
for label, pattern in rules:
text, n = re.subn(pattern, f"[{label}]", text)
if n:
counts[label] = counts.get(label, 0) + n
return text, counts
if __name__ == "__main__":
ticket = {"ticket_id": "T-1001", "name": "Ada Fictional", "birth_date": "1961-03-03",
"category": "repairs", "message": "Ada Fictional here, boiler broken. Ring 07700 900123 "
"or email ada@example.com. Postcode XX1 1XX, NI QQ 12 34 56 A."}
kept = minimise(ticket)
print("Kept fields:", list(kept))
print(*redact(kept["message"]), sep="\n")python redact.py
Kept fields: ['ticket_id', 'category', 'message']
[NAME] here, boiler broken. Ring [PHONE] or email [EMAIL]. Postcode [POSTCODE], NI [NINO].
{'NAME': 1, 'EMAIL': 1, 'PHONE': 1, 'NINO': 1, 'POSTCODE': 1}re.subn returns the new text and the number of replacements. Log such counts, never the originals, or your logs become a second copy. Now test it: save this beside it as test_redact.py.
"""test_redact.py: what does redact() catch and miss? All messages invented."""
from redact import redact
CASES = [ # (message, text that must go, text that must stay)
("Email sam@example.org or ring 020 7946 0321.", ["sam@example.org", "020 7946 0321"], []),
("Jo Testperson, NI QQ123456B, XX2 3XX.", ["Jo Testperson", "QQ123456B", "XX2 3XX"], []),
("Mobile +44 7700 900456, office 0113 496 0789.", ["7700 900456", "0113 496 0789"], []),
("My neighbour Kit Imaginary parks badly.", ["Kit Imaginary"], []),
("Ada says the boiler failed.", ["Ada"], []),
("Write to ada at example dot com.", ["ada at example dot com"], []),
("I am the only wheelchair user in Block C.", ["only wheelchair user in Block C"], []),
("Nurse, born 3 March 1961, district XX9.", ["3 March 1961", "XX9"], []),
("Router model XR7 2AB keeps rebooting.", [], ["XR7 2AB"]),
]
total = leaked = lost = 0
for i, (text, must_go, must_stay) in enumerate(CASES, 1):
out = redact(text)[0]
left = sum(s in out for s in must_go) # True counts as 1
broken = sum(s not in out for s in must_stay)
total, leaked, lost = total + len(must_go), leaked + left, lost + broken
print(i, "MISS" if left or broken else "ok ", out)
print(f"{total} to remove: {total - leaked} caught, {leaked} leaked; {lost} over-redacted")python test_redact.py
1 ok Email [EMAIL] or ring [PHONE].
2 ok [NAME], NI [NINO], [POSTCODE].
3 ok Mobile [PHONE], office [PHONE].
4 MISS My neighbour Kit Imaginary parks badly.
5 MISS Ada says the boiler failed.
6 MISS Write to ada at example dot com.
7 MISS I am the only wheelchair user in Block C.
8 MISS Nurse, born 3 March 1961, district XX9.
9 MISS Router model [POSTCODE] keeps rebooting.
13 to remove: 7 caught, 6 leaked; 1 over-redactedSeven of thirteen caught, six leaked, one product code destroyed. The cases were chosen to be hard, so this is no general miss rate: measure yours on invented test data that looks like your real traffic.
Try it yourself · Activity 04
15 minRun it and close one gap
Run both scripts, then make one improvement and measure it.
- Run
python redact.pyandpython test_redact.py, and compare with the outputs above. - Add
"Kit Imaginary"toNAMES(with a comma before it), predict the last line, then rerun the tests. Keep the change.
Could any list catch every name your users type?
Worked answer
The first runs match the outputs above. Then case 4 prints 4 ok My neighbour [NAME] parks badly. and the last line is 13 to remove: 8 caught, 5 leaked; 1 over-redacted. A list knows only names you already have; tenants mention people you have never heard of.
Keep learning
The complete workbook
This workbook is for people building or buying an AI feature in a UK organisation. You will learn where personal data hides, how the UK GDPR principles apply, when a DPIA is needed and who is controller or processor. Then you will build and honestly test a small redaction pipeline, write a retention rule and review a data flow.
- 01Where personal data hides in an AI featureIn the workbook · 1 exercise
You cannot minimise personal data until you have found it, and in an AI feature it spreads further than most teams expect.
- 02Principles, lawful basis, transparency and rightsIn the workbook · 1 exercise
Article 5 of the UK GDPR sets seven principles. Here they are, applied to one AI feature.
- 03DPIAs, roles and international transfersIn the workbook · 1 exercise
Settle three questions before personal data reaches an AI service: how risky is it, who is responsible, and where does it go?
- 04Build and test a redaction pipelineRead here · 1 exercise
Drop the fields you do not need, then redact identifiers from the free text that remains.
- 05What redaction cannot doIn the workbook · 1 exercise
The misses are the lesson: each is a kind of identifier that pattern matching handles badly.
- 06Records, retention and a data-flow reviewIn the workbook · 1 exercise
Accountability means showing your working. Finish with the records a reviewer will ask for.
Also inside: a 10-point checklist, a glossary of 12 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.
No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.
Test yourself
Questions and answers
Is this workbook legal advice?
No. It is general education based on primary sources read on 26 September 2026. Law and ICO guidance change, and many ICO guidance pages say they are under review after the Data (Use and Access) Act 2025. Involve your data protection officer early, and a solicitor for decisions about lawfulness, contracts, transfers or high risk that remains after mitigation.
Is a prompt personal data?
It is if it relates to an identified or identifiable living person, directly or indirectly. A prompt holding a name, an email address or a description that points to one person is personal data, and so are logs that copy it and any output about that person.
Are embeddings personal data?
They can be. Embeddings turn text into lists of numbers. The ICO says converting personal data into a less human-readable form, such as numbers, does not by itself mean it is no longer personal data. Treat an index built from personal data as personal data unless you can show otherwise.
If we redact prompts, is the data anonymous?
Usually not. Redaction misses unlisted names, reworded details, indirect descriptions and combinations of details. If you or anyone else can reasonably link the data back to a person, for example through a ticket number, it is at best pseudonymised, and the ICO says pseudonymised data is still personal data.
Do we need a DPIA for an AI feature?
Probably. Article 35 requires one for processing likely to result in a high risk, and the ICO's AI guidance says using AI will trigger that requirement in the vast majority of cases. If you decide you do not need one, document how you decided. Start early and ask your DPO for advice.
Is our AI provider a processor or a controller?
It depends on who decides the purposes and means. In the ICO's example, a provider answering clients' queries through an API is likely a processor for those queries but a controller for training its own model. The ICO says an organisation with no purpose of its own, acting only on a client's instructions, is likely a processor. Read the provider's terms and take advice.
What must a contract with an AI processor include?
Article 28 of the UK GDPR lists terms including: processing only on your documented instructions, confidentiality, security, conditions for using sub-processors, help with rights requests and DPIAs, deleting or returning the data at the end, and giving you the information needed to show compliance, including audits.
How long should we keep prompts and logs?
The ICO says the UK GDPR sets no specific time limits. Keep them only as long as your purpose needs, write the period down, justify it and review it. Then delete or anonymise, including backups and any copies the provider holds. The 30 days in this workbook is an invented example, not a rule.
What changed with the Data (Use and Access) Act 2025?
It amends, rather than replaces, the UK GDPR and the Data Protection Act 2018, and the ICO says all its data protection provisions are in force. Changes include a recognised legitimate interest basis, wider room for significant automated decisions with safeguards, reasonable and proportionate searches for access requests, and a duty to handle complaints. Check the ICO's current guidance, as many pages are under review.
Can we send personal data to an AI service hosted abroad?
Only within the transfer rules. If the UK GDPR applies, you initiate the transfer and the receiver is a separate organisation outside the UK, the ICO calls it a restricted transfer. It needs UK adequacy regulations, an appropriate safeguard with a transfer risk assessment, or an exception. If none applies, you must not make the transfer.
When you have finished
Get your certificate of completion
Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.
Learn the language
Key terms
- Personal data
- Information relating to an identified or identifiable living person, directly or indirectly.
- Special category data
- More sensitive personal data, such as health, ethnic origin or religious beliefs, which needs an extra condition to process.
- Controller
- The organisation that decides the purposes and means of processing personal data.
- Processor
- An organisation that processes personal data on a controller's behalf and instructions.
- Data minimisation
- Personal data should be adequate, relevant and limited to what is necessary for the purpose.
- Lawful basis
- One of the grounds in Article 6 that makes processing lawful, such as contract or legitimate interests.
6 of the workbook's 12 terms. The complete glossary is in the workbook.
Follow the evidence
Sources and checks
Facts last checked: .
Examples in this workbook were run on: Python 3.12.10 on Windows 11 (standard library only) (2026-09-26).
These workbooks use AI assistance. See how the workbooks are made.
- Guidance on AI and data protection (under review after the Data (Use and Access) Act; read 26 September 2026)Information Commissioner's Office (ICO)
- What are the accountability and governance implications of AI? (DPIAs and controller or processor roles in AI)Information Commissioner's Office (ICO)
- How do we ensure lawfulness in AI?Information Commissioner's Office (ICO)
- How do we ensure transparency in AI?Information Commissioner's Office (ICO)
- How should we assess security and data minimisation in AI?Information Commissioner's Office (ICO)
- How do we ensure individual rights in our AI systems?Information Commissioner's Office (ICO)
- A guide to the data protection principlesInformation Commissioner's Office (ICO)
- Principle (c): Data minimisationInformation Commissioner's Office (ICO)
- Principle (e): Storage limitationInformation Commissioner's Office (ICO)
- What is personal data?Information Commissioner's Office (ICO)
- A guide to lawful basis (updated 2 April 2026)Information Commissioner's Office (ICO)
- Right to be informedInformation Commissioner's Office (ICO)
- Data Protection Impact Assessments (DPIAs)Information Commissioner's Office (ICO)
- When do we need to do a DPIA?Information Commissioner's Office (ICO)
- Data protection officersInformation Commissioner's Office (ICO)
- International transfersInformation Commissioner's Office (ICO)
- A brief guide to international transfers (15 January 2026)Information Commissioner's Office (ICO)
- Introduction to anonymisationInformation Commissioner's Office (ICO)
- How do we ensure anonymisation is effective?Information Commissioner's Office (ICO)
- Who needs to document their processing activities?Information Commissioner's Office (ICO)
- What do we need to document under Article 30 of the UK GDPR?Information Commissioner's Office (ICO)
- Data (Use and Access) Act 2025Information Commissioner's Office (ICO)
- The Data Use and Access Act 2025 (DUAA): what does it mean for organisations? (updated 19 June 2026)Information Commissioner's Office (ICO)
- The Data Use and Access Act 2025 (DUAA): summary of the changes to data protection lawInformation Commissioner's Office (ICO)
- Data Use and Access Act 2025: plans for commencement (last updated 5 February 2026)Department for Science, Innovation and Technology and DCMS (GOV.UK)
- Data (Use and Access) Act 2025 (c. 18)legislation.gov.uk (The National Archives)
- UK GDPR, Article 4: Definitionslegislation.gov.uk (The National Archives)
- UK GDPR, Article 5: Principles relating to processing of personal datalegislation.gov.uk (The National Archives)
- UK GDPR, Article 6: Lawfulness of processinglegislation.gov.uk (The National Archives)
- UK GDPR, Article 25: Data protection by design and by defaultlegislation.gov.uk (The National Archives)
- UK GDPR, Article 28: Processorlegislation.gov.uk (The National Archives)
- UK GDPR, Article 30: Records of processing activitieslegislation.gov.uk (The National Archives)
- UK GDPR, Article 35: Data protection impact assessmentlegislation.gov.uk (The National Archives)
- Data Protection Act 2018, section 171: Re-identification of de-identified personal datalegislation.gov.uk (The National Archives)
- Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (18 December 2024)European Data Protection Board (EDPB)
- Guidelines for secure AI system developmentNational Cyber Security Centre (NCSC)
- Telephone numbers for use in TV and radio drama programmesOfcom
- Example DomainsInternet Assigned Numbers Authority (IANA)
- NIM39110: National Insurance Numbers (NINOs): what a NINO looks likeHM Revenue and Customs (GOV.UK)
- re: Regular expression operationsPython Software Foundation
Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 26 September 2026.
NextKeep going
Where to go next
Recommended for you
Host a chatbot on your website, safely
Put a chatbot on your site without leaking keys or running up bills: server function, limits, CORS, privacy, accessibility, fallbacks, monitoring and a launch checklist.
Recommended for you
RAG and knowledge bases, step by step
Build a knowledge base a model can answer from: collect, clean, chunk, embed, retrieve, cite and evaluate, with security basics and a hands-on exercise using five short documents.
Recommended for you
How to evaluate a new frontier model the day it lands
A ten-step method for judging a new model: primary sources, licence, architecture, hardware, benchmarks, your own tests, safety, jurisdiction and a decision record.