# Automatically anonymizing personal data with AI

[Skip to content](#lm-inhoud)Network/[NL](/en/gevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai)EN[Hubhub.llmnet.nlCompare models on task, language, cost and license.](https://hub.llmnet.nl/en/)[Communitycommunity.llmnet.nlPrompt techniques, patterns and system prompts.](https://community.llmnet.nl/en/)[APIapi.llmnet.nlLLMs in production: rate limits, routing, structured output.](https://api.llmnet.nl/en/)[Consultancyconsultancy.llmnet.nlRolling out AI in an organization, pilot to production.](https://consultancy.llmnet.nl/en/)[Newsnieuws.llmnet.nlAI developments, explained for the Netherlands.](https://nieuws.llmnet.nl/en/)[Benchmarkbenchmark.llmnet.nlMeasure AI quality yourself, on your own tasks.](https://benchmark.llmnet.nl/en/)[Careersvacatures.llmnet.nlAI roles, salaries and career paths in the Netherlands.](https://vacatures.llmnet.nl/en/)[Learnleren.llmnet.nlAI concepts in plain language, beginner to builder.](https://leren.llmnet.nl/en/)[Guidegids.llmnet.nlRun AI privately on your own Mac, PC, NAS or home server.](https://gids.llmnet.nl/en/)[Directorydirectory.llmnet.nlMapping the AI ecosystem: tools, models, companies.](https://directory.llmnet.nl/en/)[Radarradar.llmnet.nlSignals from X, research and communities for indie developers.](https://radar.llmnet.nl/en/)[Appsapps.llmnet.nlReviews of AI apps and open-source repos, with tips for builders.](https://apps.llmnet.nl/en/)[llmnet.nl — main site](https://llmnet.nl/en/)[](https://x.com/intent/post?url=https%3A%2F%2Fgids.llmnet.nl%2Fen%2Fgevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai&text=Automatically%20anonymizing%20personal%20data%20with%20AI)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fgids.llmnet.nl%2Fen%2Fgevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai)[](https://www.reddit.com/submit?url=https%3A%2F%2Fgids.llmnet.nl%2Fen%2Fgevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai&title=Automatically%20anonymizing%20personal%20data%20with%20AI)[](#)[](https://x.com/intent/post?url=https%3A%2F%2Fgids.llmnet.nl%2Fen%2Fgevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai&text=Automatically%20anonymizing%20personal%20data%20with%20AI)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fgids.llmnet.nl%2Fen%2Fgevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai)[](https://www.reddit.com/submit?url=https%3A%2F%2Fgids.llmnet.nl%2Fen%2Fgevoelige-persoonsgegevens-automatisch-anonimiseren-met-ai&title=Automatically%20anonymizing%20personal%20data%20with%20AI)[](#)

 
# Automatically anonymizing personal data with AI

 By Ivo Donker — compiled with AI assistance (Claude & Gemini)

 Within the path of selecting, setting up, and operationally managing local language models, privacy protection is the crucial link between raw data and safe automation. As soon as internal customer contacts, Woo requests (Dutch Open Government Act), medical notes, or legal files are processed, there's a real risk that directly identifiable personal data ends up unintentionally in context windows, vector indices, or fine-tuned models. To precisely match the required processing power for local anonymization tasks to the model size, the reference overview on [hardware for local LLMs](https://gids.llmnet.nl/en/hardware-voor-lokale-llm) offers insight into the distribution between CPU cores, system RAM, and GPU memory.

 Manually redacting documents is time-consuming, costly, and inherently prone to human fatigue. On the other hand, traditional regex searches fall short as soon as personal data is interwoven with everyday language or indirect context. By building a layered architecture in which deterministic pattern recognition, Named Entity Recognition (NER), and local neural language models check each other, a robust processing pipeline emerges. In the process, not a single unfiltered byte leaves your own server infrastructure.

 
## The spectrum under the GDPR: from redacting to synthetic replacement

 Under the General Data Protection Regulation (GDPR), there's a sharp legal and technical dividing line between pseudonymization and irreversible anonymization. With pseudonymization, direct identifiers — such as a first and last name or a Dutch citizen service number (BSN) — are replaced with a unique token, hash, or pseudonym (for example KLANT_8492). Because a mapping key exists somewhere that can restore the original identity, the data legally remains personal data. The GDPR continues to apply in full to the entire dataset, including the obligations around processing agreements, data breach notifications, and retention periods.

 With true anonymization, tracing back to a natural person has been made technically irreversible using reasonably available means. To map out the organizational conditions and formal privacy obligations, the [GDPR privacy checklist for organizations](https://gids.llmnet.nl/en/avg-privacy-checklist) helps verify which processing steps are necessary within local workflows.

 
 
 
 
 Method | 
 Mechanism | 
 Reversibility | 
 Semantic Preservation | 
 GDPR Qualification | 
 

 
 
 
 Redacting (Blackout) | 
 Replaced by fixed blocks such as [GEBLOKKEERD] | 
 Irreversible | 
 Very low; breaks grammatical sentence structure | 
 Anonymous (provided no indirect data remains) | 
 

 
 Categorical Masking | 
 Replaced by type tags such as <DATUM> or <BSN> | 
 Irreversible | 
 Moderate; logic remains partly intact, context becomes fragmented | 
 Anonymous (in the absence of quasi-identifiers) | 
 

 
 Classic Pseudonymization | 
 Replaced by a consistent hash with a secret salt | 
 Reversible via key table | 
 Good for aggregations, moderate for language models | 
 Personal data (GDPR still applies) | 
 

 
 Synthetic Replacement | 
 Replaced by realistic fictional names and locations | 
 Irreversible (fiction without a key) | 
 Excellent; preserves fluent syntax and embeddings | 
 Anonymous (with correct differentiation) | 
 

 
 
 

 
## Why regular expressions and keyword lists fail

 Deterministic filters based on regular expressions (regex) are fast, use almost no compute power, and work flawlessly for structured data formats. A Dutch account number (IBAN), a license plate, or a BSN that passes the mathematical eleven-test can be extracted via regex with nearly one hundred percent certainty. The problem arises, however, as soon as personal data is embedded in natural, unstructured language.

 A typical Dutch sentence illustrates the shortcomings of static rules: We hebben gisteren met De Graaf overlegd in Den Bosch over de renovatie van De Zwaan. (lit.: "We met with De Graaf yesterday in Den Bosch to discuss the renovation of De Zwaan.") A simple filter sees 'De Graaf' as a noble title or common noun, 'Den Bosch' as a place name (or surname), and 'De Zwaan' as a bird, a pub, or a surname. Without semantic insight, two problems arise:

 1. False Negatives: Names that coincide with common nouns (Bakker, Visser, De Boer, Koster) or rare foreign names are overlooked. This directly leads to a data breach.
 2. False Positives: Words such as 'Monday', 'Director', or 'Main Street' get masked incorrectly, rendering the document's actual meaning unusable.

 In addition, documents contain quasi-identifiers: data that by itself contains no direct name but, in combination, uniquely identifies an individual. Think of: the only female neurologist at the regional hospital in Boxmeer who earned her doctorate in 2021. No regular expression detects this sentence as personal data, while the person involved can be identified with a single search.

 
## The layered architecture: Presidio, spaCy, and local neural models

 To catch both structured identifiers and deep contextual traces without crippling processing speed, we build a three-layer inspection pipeline. Each layer has a specific purpose, its own latency profile, and a well-defined task.

 The first layer is the deterministic pre-filter. This runs Microsoft Presidio Analyzer with custom recognizers for Dutch standards: BSN, IBAN, phone numbers, email addresses, and postal codes. This layer catches 60 to 70 percent of raw identifiers in under two milliseconds per page.

 The second layer is the fast statistical NER layer. Here, a compact transformer model or spaCy pipeline (such as nl_core_news_lg or a fine-tuned Dutch RoBERTa model) analyzes the grammatical sentence structure. The model marks personal names, organizations, locations, and dates based on their syntactic role in the sentence. This takes about 15 to 40 milliseconds per page on a modern processor.

 The third layer is the semantic LLM evaluator. Only text sections with a low confidence score, complex indirect descriptions, or heavily intertwined job titles are passed to a local language model (such as Llama 3 or Qwen 2.5). This model recognizes implicit connections and indirect identifiers missed by the earlier layers.

 To understand how these local components run safely within the corporate network without dependence on external cloud providers, the article on [privacy-friendly AI use](https://gids.llmnet.nl/en/privacyvriendelijk-ai) explains which network isolation and data storage rules are needed. For a broader overview of the legislative frameworks around model choices, the dossier on [AI models and GDPR compliance](https://hub.llmnet.nl/en/ai-modellen-en-privacy-avg-compliance) offers further guidelines.

 
## Implementation: a working Python pipeline with Microsoft Presidio

 The implementation below demonstrates setting up a local anonymization engine in Python. The code combines Presidio with a custom validator for the Dutch citizen service number, including the mandatory eleven-test check to rule out false 9-digit sequences.

 from presidio_analyzer import AnalyzerEngine, PatternRecognizer, Pattern
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig

# Stap 1: BSN-patroon definiëren (9 aaneengesloten cijfers of met scheidingstekens)
bsn_patroon = Pattern(
 name="nl_bsn_regex",
 regex=r"\b\d{9}\b",
 score=0.70
)

# Stap 2: Aangepaste herkenner met elfproef-validatie
class NederlandseBsnRecognizer(PatternRecognizer):
 def __init__(self):
 super().__init__(
 supported_entity="NL_BSN",
 patterns=[bsn_patroon],
 supported_language="nl"
 )

 def validate_result(self, pattern_text: str) -> bool:
 schone_tekst = pattern_text.strip()
 if len(schone_tekst) != 9 or not schone_tekst.isdigit():
 return False
 
 cijfers = [int(c) for c in schone_tekst]
 # De officiële 11-proef voor BSN: 9*c1 + 8*c2 + ... + 2*c8 - 1*c9
 som = sum(cijfers[i] * (9 - i) for i in range(8)) - cijfers[8]
 return som % 11 == 0 and som != 0

# Stap 3: Initialiseer de engine met Nederlandse en Engelse ondersteuning
analyzer = AnalyzerEngine(supported_languages=["nl", "en"])
analyzer.registry.add_recognizer(NederlandseBsnRecognizer())
anonymizer = AnonymizerEngine()

# Testtekst met gemengde entiteiten
brondocument = (
 "Betrokkene J. de Vries (BSN: 111222333) meldde op 14 augustus dat er "
 "onregelmatigheden zijn geconstateerd bij vestiging Alkmaar via info@bedrijf.nl."
)

# Detecteer gevoelige entiteiten
resultaten = analyzer.analyze(
 text=brondocument,
 language="nl",
 entities=["NL_BSN", "EMAIL_ADDRESS", "PERSON", "LOCATION"]
)

# Pas categorische maskering toe
gemaskeerde_uitvoer = anonymizer.anonymize(
 text=brondocument,
 analyzer_results=resultaten,
 operators={
 "NL_BSN": OperatorConfig("replace", {"new_value": "<BSN_GEANONIMISEERD>"}),
 "EMAIL_ADDRESS": OperatorConfig("replace", {"new_value": "<EMAIL_GEANONIMISEERD>"}),
 "PERSON": OperatorConfig("replace", {"new_value": "<PERSOON>"}),
 "DEFAULT": OperatorConfig("replace", {"new_value": "<VERTROUWELIJK>"})
 }
)

print(gemaskeerde_uitvoer.text)

 
## Contextual entity detection and quasi-identifiers via local LLMs

 When texts contain complex phrasing, a static library such as Presidio sometimes falls short. We can deploy a local model (for example via Ollama or vLLM) as a second inspection round. Here, we force the model to return only a strict JSON schema with the detected offsets and entity types.

 To ensure the local model communicates without parsing errors, it's advisable to consult the guide on [enforcing structured JSON outputs in local language models](https://gids.llmnet.nl/en/structured-outputs-lokale-llm) . This ensures the model always produces valid coordinates and entities that can be processed programmatically right away.

 The prompt instruction forces the model to look beyond simple proper nouns and explicitly search for relationships, rare job titles, and indirect characteristics:

 {
 "entities": [
 {
 "start": 14,
 "end": 42,
 "type": "INDIRECT_IDENTIFIER",
 "text": "enig overgebleven cardioloog",
 "reason": "Beroep in combinatie met kleine afdeling maakt persoon uniek traceerbaar"
 },
 {
 "start": 68,
 "end": 89,
 "type": "QUASI_IDENTIFIER",
 "text": "geboren op 29-02-1968",
 "reason": "Zeldzame geboortedatum (schrikkeldag) verhoogt herleidbaarheid"
 }
 ]
}

 
## Synthetic data injection: preserving semantics and embedding quality

 Traditionally redacting or replacing names with static tags such as <PERSOON_1> or [ONBEKEND] seriously damages the internal syntax and semantics of texts. When such documents are then loaded into a vector search engine or RAG architecture (Retrieval-Augmented Generation), embedding models perform significantly worse. A vector representation of <PERSON> visited <LOCATION> because of <CONDITION> lacks the contextual nuances needed for accurate semantic similarity calculations.

 Synthetic replacement solves this structurally. Instead of mutilating the text, the system generates consistent, contextually appropriate fictional alternatives. A Dutch name like 'Jan Willem van den Berg' is replaced with 'Pieter Schipper', an address in Groningen is replaced with a non-existent address in Zwolle, and an industry sector is preserved without revealing the actual trade name. This process guarantees:

 1. Grammatical Integrity: Articles, prepositions, and verb conjugations remain syntactically correct.
 2. Consistent Entity Linking: If a person appears multiple times in a document, exactly the same synthetic pseudonym is used everywhere.
 3. Preserved Embedding Distances: Vectors of synthetic documents cluster in vector databases in a similar way to the original text.

 For anyone planning to make anonymized text files locally indexable, the guide on [searching documents with local RAG](https://gids.llmnet.nl/en/documenten-doorzoeken-lokaal) explains how anonymized source files are optimally tokenized and converted into embeddings.

 
## Measurement methodology and quality assurance: recall, precision, and F2 score

 When measuring anonymization quality, we apply fundamentally different acceptance criteria than for general Natural Language Processing tasks. A classification error is not evenly weighted here: a missed piece of personal data (False Negative) constitutes a potential data breach and a GDPR violation, while an incorrectly masked neutral word (False Positive) merely results in minor textual noise.

 
 
 
 
 Metric | 
 Mathematical Definition | 
 Production Target | 
 Practical Meaning | 
 

 
 
 
 Recall (Sensitivity) | 
 TP / (TP + FN) | 
 > 99.5% | 
 What percentage of all actual personal data was successfully detected? | 
 

 
 Precision | 
 TP / (TP + FP) | 
 > 93.0% | 
 What percentage of the masked words was actually personal data? | 
 

 
 F2 Score | 
 5 · (P · R) / (4 · P + R) | 
 > 0.98 | 
 Weighted harmonic mean in which Recall weighs four times as heavily as Precision. | 
 

 
 
 

 To reliably establish these scores in a production environment, it's necessary to build a gold standard (gold standard dataset). This set should consist of at least 300 manually annotated Dutch documents from your own domain. Every update to a regex pattern, spaCy version, or neural model must be automatically validated against this dataset to immediately detect regressions.

 
## Domain-specific edge cases: medical, legal, and financial

 In regulated sectors, specific linguistic and statistical patterns occur that bypass standard anonymization tools. Without domain-specific rules, the risk of de-anonymization remains unacceptably high.

 
### Medical Documentation

 Electronic patient records (EPRs) contain rare diagnoses that, combined with basic demographic data, directly make a patient unique. According to the principle of k-anonymity , a record within a dataset must be indistinguishable from at least $k - 1$ other individuals. A diagnosis such as fibrodysplasia ossificans progressiva in a 14-year-old boy in Friesland directly identifies one specific individual in the Netherlands. The anonymization pipeline must generalize such rare conditions to a higher ICD-10 category (for example "rare connective tissue disorder").

 
### Legal documents and Woo decisions

 In legal rulings and disclosures under the Dutch Open Government Act, not only the names of the parties involved must be removed, but also case numbers, docket numbers, cadastral references, and incident timestamps. A case number such as NL24.18492 leads, via public case-law registers, to the full personal data of the parties involved within seconds.

 
### Financial Transactions

 Besides IBAN numbers, financial descriptions often contain payment references, order numbers, Chamber of Commerce numbers, and specific transaction amounts with decimals (for example "invoice 2024-8841 in the amount of €14,821.37"). Because exact amounts can be correlated with annual accounts or bank statements, amounts should be rounded to orders of magnitude (for example "between €10,000 and €20,000") or replaced synthetically.

 
## Throughput speeds, hardware requirements, and operational costs

 The choice of processing depth has direct consequences for the required hardware and the system's throughput capacity. Running a local neural language model requires significantly more compute power than traditional pattern recognition.

 
 
 
 
 Pipeline Configuration | 
 Minimum Hardware | 
 Throughput | 
 Estimated Processing Time (10,000 docs.) | 
 

 
 
 
 Pure Regex + Presidio | 
 4 CPU cores, 4 GB RAM | 
 ~350 pages/sec | 
 less than 1 minute | 
 

 
 Regex + spaCy NER (CPU) | 
 8 CPU cores, 8 GB RAM | 
 ~45 pages/sec | 
 about 4 minutes | 
 

 
 Hybrid: Regex + spaCy + Llama-3-8B (GPU) | 
 1x RTX 4090 (24 GB VRAM) or M2/M3 Max (36 GB) | 
 ~12 pages/sec | 
 about 14 minutes | 
 

 
 Full LLM Analysis per Document | 
 2x RTX 4090 or A100 (80 GB) | 
 ~2 pages/sec | 
 about 1.4 hours | 
 

 
 
 

 In practice, the hybrid structure offers the best balance between cost and accuracy: 90 percent of the text is handled lightning-fast by CPU-based pattern and NER models, and only the remaining 10 percent of borderline cases is forwarded to the GPU-based LLM evaluator.

 
## Step-by-step operational checklist for implementation

 To roll out a reliable, GDPR-compliant anonymization pipeline within the organization, the implementation team goes through the following steps:

 1. Data Inventory and Entity Definition: Determine exactly which data categories occur in the source files (BSN, name/address/city, BIG registration numbers, license plates, salary data).
 2. Building a Validation Corpus: Build a manually annotated test set of at least 250 representative documents with varying writing styles.
 3. Regex and Checksum Configuration: Implement deterministic validators with mathematical checks (such as the 11-test for BSN and IBAN modulo validations).
 4. NER Model Selection and Fine-Tuning: Connect a Dutch transformer model and test whether typical Dutch surname prefixes (such as 'van der', 'de', 'in 't') are correctly attached to surnames.
 5. Setting Up a Structured LLM Fallback: Connect a local LLM with a strict JSON schema contract for detecting quasi-identifiers in complex paragraphs.
 6. Choosing the Replacement Strategy: Determine per downstream system whether categorical masking (Woo publications), pseudonymization with a key table, or synthetic reconstruction (RAG/embeddings) is required.
 7. Audit Logging Without Data Leaks: Log processing counts, entity types, and confidence scores in monitoring logs, but never store the original sensitive text fragments in them.
 8. Periodic Quality Audit: Run a regression test on the validation corpus every month to detect model drift and changes in language conventions in time.

 This methodical setup creates a privacy-friendly processing pipeline with which organizations can fully automatically process, analyze, and unlock large amounts of sensitive information, without making concessions on data breach prevention or legal compliance.
