Messy Real-World Text

Finding and masking personal data in text

PII detection finds personal data hiding inside free text, like an email or phone number typed into a support message, and masks it before the text goes anywhere else.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

PII detection finds personal information hiding inside ordinary text. An email in a complaint, a phone number in a comment — both get masked before the text goes anywhere.

Think about a photocopier operator handling a stack of documents for a public report. Before copying, a careful operator blacks out account numbers and signatures on each page. The report still makes sense. The private details do not travel with it.

PII detection is that same blacking-out step, done automatically on text instead of paper. PII stands for personally identifiable information — anything that could identify a specific real person.

Why it exists

Support tickets, reviews and chat logs constantly contain personal data typed in passing. "Email me at ravi@example.com" might sit inside a message really about a delivery delay. That message might get logged, shared with an analytics team, or fed into a model for training.

Privacy law and basic good practice both require limiting who sees that personal data, and for how long. Manually reading every message to redact sensitive details does not scale. Automated PII detection finds the patterns and masks them before the text moves further downstream.

How it works

  "Contact me at ravi@example.com or 98765 43210"
                      |                    |
                      v                    v
                  [EMAIL]              [PHONE]

Some kinds of PII — emails, phone numbers, ID number formats — follow predictable patterns that a rule can match directly. Others — a person's name, a home address — need a model that understands context. Names do not follow one fixed shape.

Where you have already seen it

  • Customer support platforms automatically masking card numbers and emails in ticket logs before an agent sees them.
  • Data anonymisation before a dataset is shared with a research team or used to train a model.
  • Chat moderation tools flagging when a user is about to publicly post their own phone number.

Remember this

  • PII detection finds personal data that appears inside otherwise-unstructured text, not in a dedicated, labelled field.
  • Pattern-shaped PII — emails, phone numbers, ID formats — can be caught with regular expressions directly.
  • Free-form PII — names, addresses — needs a context-aware model, since no fixed pattern describes every name.

What to learn next

Developer — Code and libraries.

Setup

Nothing to install. Pure Python standard library — re.

Detecting and masking common PII patterns

pii_redact_demo.py
import re

text = """
Contact Rahul at rahul.sharma@example.com or call +91 98765 43210.
His Aadhaar number is 1234 5678 9012 and PAN is ABCDE1234F.
Backup email: r.sharma123@company.co.in
"""

PATTERNS = {
    "EMAIL": r"[\w.+-]+@[\w-]+\.[\w.-]+",
    "PHONE_IN": r"(?:\+91[\s-]?)?[6-9]\d{4}[\s-]?\d{5}",
    "AADHAAR": r"\b\d{4}\s\d{4}\s\d{4}\b",
    "PAN": r"\b[A-Z]{5}\d{4}[A-Z]\b",
}

def redact(text, patterns):
    for label, pattern in patterns.items():
        text = re.sub(pattern, f"[{label}]", text)
    return text

print(redact(text, PATTERNS))
Output

Contact Rahul at [EMAIL] or call [PHONE_IN].
His Aadhaar number is [AADHAAR] and PAN is [PAN].
Backup email: [EMAIL]

Line by line

Every pattern targets one specific, well-defined PII shape. EMAIL matches the general local-part-at-domain structure. PHONE_IN matches a 10-digit Indian mobile number, optionally preceded by the country code, starting with a digit from 6 to 9 — the actual valid range for Indian mobile numbers. AADHAAR and PAN match India's national ID number formats exactly.

Patterns are applied in sequence, each replacing its own match with a label. Order rarely matters here, since the patterns target visually distinct shapes that do not overlap with each other in this example.

Both email addresses got caught, including the one with digits and a country-code-style domain. r.sharma123@company.co.in matches the general email pattern correctly, even though it looks structurally different from the first email.

Common mistakes

Assuming a regex-based approach catches every kind of PII. Names, physical addresses and other free-form personal data have no single fixed shape a regex can reliably describe — those need a context-aware model, typically the named-entity-recognition approach covered in named entity recognition.

Writing a phone number pattern that is too loose. A pattern like \d{10} matches any ten consecutive digits — including a ten-digit order number or invoice ID that is not a phone number at all, producing false positives that redact things that were never sensitive.

Redacting in a way that destroys useful structure. Replacing every match with the exact same generic [REDACTED] loses information a downstream analytics process might legitimately need, like "this message contained a phone number" versus "this message contained an email." Labelled redaction, as shown above, keeps that structure while still hiding the actual value.

Try it yourself

Add a false-positive test: a ten-digit order number that is not a real Aadhaar or phone number, and see whether any of the current patterns wrongly match it. Then tighten the relevant pattern — using word boundaries, or context words like "order" or "invoice" nearby — to reduce that false match without breaking the correct matches shown above.

What to learn next

Researcher — Mathematics and papers.

Two detection paradigms

Pattern-based detection, as shown in the developer demo, matches PII with a fixed structural shape — emails, phone numbers, national ID formats, credit card numbers (which additionally validate against the Luhn checksum algorithm) — with high precision and near-zero training cost, at the price of zero recall on anything outside the hand-specified pattern set.

Model-based detection, typically framed as named entity recognition with PII-specific entity types (PERSON, LOCATION, ORGANIZATION, and finer-grained types like MEDICAL_RECORD_NUMBER in domain-specific systems), catches free-form PII with no fixed shape, at the cost of needing labelled training data and inherently probabilistic, imperfect recall — a name genuinely can be missed, or a non-name can be flagged, unlike a well-specified regex.

Production PII detection systems, such as Microsoft Presidio and AWS Comprehend PII, combine both: pattern-based recognizers for structured PII, model-based recognizers for names and addresses, and a confidence-scoring layer that merges and deduplicates overlapping detections from each.

Formalising the precision-recall trade-off

For a PII detector, precision and recall carry asymmetric real costs:

text
precision = true_positives / (true_positives + false_positives)
recall    = true_positives / (true_positives + false_negatives)
  • A false positive (over-redaction) degrades downstream usefulness of the text — costly, but recoverable and safe.
  • A false negative (under-redaction) is a genuine privacy or compliance failure — a real person's data has actually leaked through undetected.

This asymmetry is why production PII systems are commonly tuned to favour recall over precision, accepting more over-redaction as the safer failure mode, in contrast to many other NLP classification tasks where the two error types are treated as roughly equally costly.

Regulatory context

The specific PII categories a system must detect, and the retention and consent rules governing them, are set by jurisdiction-specific law rather than by any technical standard: India's Digital Personal Data Protection Act (2023), the EU's GDPR, and sector-specific rules (health data, financial data) each define personal data somewhat differently and impose different obligations once it is detected. PII detection tooling implements the technical capability; the legal category boundaries it must respect are set externally, not by the NLP field.

Key references

  • Dernoncourt, F., Lee, J., Uzuner, O. & Szolovits, P. (2017). De-identification of Patient Notes with Recurrent Neural Networks. JAMIA. — sequence-labelling PII detection in the high-stakes clinical-text setting.
  • Lison, P., Pilán, I., Sanchez, D., Batet, M. & Øvrelid, L. (2021). Anonymisation Models for Text Data: State of the Art, Challenges and Future Directions. ACL. — a comprehensive survey of the field, covering both pattern- and model-based approaches.
  • Microsoft. Presidio: Data Protection and De-identification SDK. microsoft.github.io/presidio — a widely-deployed open-source implementation combining both detection paradigms.

Current state and open problems

Contextual PII — a piece of information harmless in isolation but identifying in combination, like "the only cardiologist who joined this hospital in March" — falls entirely outside what any per-span detector, pattern-based or model-based, can catch, since the identifying power comes from combining multiple facts across a document or dataset rather than from any single span of text. This is a structural limitation of the entire "detect and mask" approach, addressed instead by higher-level anonymisation techniques like k-anonymity, applied at the dataset level rather than the sentence level.

Cross-lingual and script-specific PII patterns — an ID number format specific to one country, a name in a script the model was not well trained on — remain a genuine, current gap, tracking the same uneven-coverage pattern already documented for named entity recognition and other model-based tasks across less-resourced languages.

What to learn next