AI Safety and Ethics

Privacy in machine learning

A trained model can leak the people it was trained on — and both the leak and the fix are measurable in a few lines of code.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why this is a real problem, not a theoretical one
  4. The three ways a model leaks
  5. The picture
  6. The good news, which is genuinely good
  7. What is honestly hard here
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A model remembers some of the people it learned from, and a stranger can sometimes tell who they were.

The analogy you have already lived

Think of a friend who read your diary once, years ago. Ask them about your life and they mostly speak in vague generalities.

But say a very specific sentence from that diary and watch their face. The flicker of recognition tells you they have seen it before. They never quoted a line. They gave it away by reacting differently to something familiar.

A model does the same flicker. It answers with more confidence about the exact examples it was trained on. That confidence gap is a leak, and it can be measured.

Why this is a real problem, not a theoretical one

Deleting names does not make data anonymous. This is the single most expensive misunderstanding in the field.

Suppose a hospital publishes patient records with the names removed. Each row still has a birth year, a pin code and a gender. In a city of millions, that combination alone identifies most people uniquely.

Anyone holding a voter roll can match the rows back to names. The diagnosis was the only thing hidden, and now it is not.

The developer block measures this on a fake table of 5,000 people. More than 60% of the rows are unique on those three harmless-looking columns.

The three ways a model leaks

It tells you who was in the training data. By answering more confidently on training examples. This is called membership inference — working out whether a specific person's record was used.

It repeats things word for word. A model trained on emails can be prompted into reproducing an address or a phone number it saw during training.

It lets you reconstruct a face or a record. From enough queries, an attacker can rebuild an approximation of an individual training example.

The picture

   real people  →  training data  →  model  →  answers
        ▲                                        │
        └────────── attacker works backwards ────┘

   "Was Ramesh's record used to train this?"
   "What phone number follows this name?"

The arrow going backwards is the whole problem. Most teams design only the forward arrow.

The good news, which is genuinely good

The leak is caused by memorising. A model that memorises its training set is both worse at its job and leakier.

So the same discipline that stops overfitting — a model memorising its training examples instead of learning the pattern — also reduces the leak. In the experiment below, one settings change drops the attack's success from 93% to 55%, and costs one percentage point of accuracy. That is an unusually cheap trade.

What is honestly hard here

There is no switch labelled "private". Every real defence costs some accuracy, and the amount is not knowable in advance.

And privacy is not only a model problem. Logs, backups, prompts sent to a third-party API, and screenshots in a support ticket all leak. The model is often the safest part of the system.

Remember this

  • Deleting names does not anonymise data — combinations of ordinary columns identify people.
  • Models leak by being more confident about their training examples.
  • Fighting memorisation is the same work as fighting overfitting, and it is cheap.

What to learn next

Developer — Code and libraries.

Measure the leak before you argue about it

Privacy in ML becomes an engineering problem the moment you can put a number on it. The number is membership inference attack AUC: how well can an attacker tell members of the training set from non-members, using only model outputs?

An AUC of 0.5 means the attacker learns nothing. An AUC of 0.9 means they are nearly always right.

Setup

bash
pip install numpy scikit-learn pandas

Runs on CPU in a couple of seconds.

Part 1 — a membership inference attack in fifteen lines

membership_inference.py
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(21)
n, d = 1000, 20
X = rng.normal(0, 1, (n, d))
y = (X[:, :3].sum(1) + rng.normal(0, 2.0, n) > 0).astype(int)   # deliberately noisy labels
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)


def attack(model, label):
    model.fit(Xtr, ytr)
    # the attacker only needs the probability the model gives to the TRUE label
    p_member = model.predict_proba(Xtr)[np.arange(len(ytr)), ytr]
    p_outsider = model.predict_proba(Xte)[np.arange(len(yte)), yte]
    scores = np.r_[p_member, p_outsider]
    is_member = np.r_[np.ones(len(ytr)), np.zeros(len(yte))]
    print(f"{label:26s} train_acc={model.score(Xtr, ytr):.3f} "
          f"test_acc={model.score(Xte, yte):.3f}  "
          f"mean_conf_member={p_member.mean():.3f} "
          f"mean_conf_outsider={p_outsider.mean():.3f}  "
          f"ATTACK_AUC={roc_auc_score(is_member, scores):.3f}")


print("An attack AUC of 0.500 means the attacker learns nothing.\n")
attack(RandomForestClassifier(n_estimators=200, random_state=0), "memorising model")
attack(RandomForestClassifier(n_estimators=200, min_samples_leaf=40, random_state=0),
       "regularised model")
Output
An attack AUC of 0.500 means the attacker learns nothing.

memorising model           train_acc=1.000 test_acc=0.744  mean_conf_member=0.852 mean_conf_outsider=0.603  ATTACK_AUC=0.932
regularised model          train_acc=0.784 test_acc=0.734  mean_conf_member=0.582 mean_conf_outsider=0.562  ATTACK_AUC=0.554

What that result actually says

The first model gives away its training set. Attack AUC 0.932. Note the two confidence columns: 0.852 for members against 0.603 for outsiders. The model is not quoting anybody's data. It is reacting differently, and that is enough.

The gap between train and test accuracy is the leak. 1.000 against 0.744 is a 25-point generalisation gap, and the attack AUC tracks it closely. Overfitting and privacy leakage are the same phenomenon viewed from two sides.

One hyperparameter closed most of the hole. min_samples_leaf=40 forces every leaf to describe at least 40 people, so no leaf is about one person. Attack AUC falls to 0.554. Test accuracy falls from 0.744 to 0.734.

But 0.554 is not 0.500. Regularisation reduces the leak; it does not bound it. For a bound you need differential privacy, which is the next lesson.

Part 2 — "anonymised" data, measured

k_anonymity.py
import numpy as np, pandas as pd

rng = np.random.default_rng(2)
N = 5000
df = pd.DataFrame({
    "name_removed": ["<redacted>"] * N,                      # the "anonymisation" people trust
    "birth_year":   rng.integers(1955, 2006, N),
    "pin_code":     rng.integers(400001, 400105, N),         # ~104 Mumbai pin codes
    "gender":       rng.choice(["F", "M"], N),
    "diagnosis":    rng.choice(["asthma", "diabetes", "none"], N),
})

quasi = ["birth_year", "pin_code", "gender"]
sizes = df.groupby(quasi).size()

print(f"rows: {N}   quasi-identifiers: {quasi}")
print(f"k-anonymity of this table (smallest group): {sizes.min()}")
print(f"rows that are UNIQUE on those three columns: {(sizes == 1).sum()} "
      f"({(sizes == 1).sum() / N:.1%})")

# Generalise: decade instead of year, first 4 digits of the pin code
df["decade"] = (df.birth_year // 10) * 10
df["pin_area"] = df.pin_code // 100
sizes2 = df.groupby(["decade", "pin_area", "gender"]).size()
print(f"\nafter generalising to decade + pin area:")
print(f"k-anonymity: {sizes2.min()}   unique rows: {(sizes2 == 1).sum()}")
Output
rows: 5000   quasi-identifiers: ['birth_year', 'pin_code', 'gender']
k-anonymity of this table (smallest group): 1
rows that are UNIQUE on those three columns: 3087 (61.7%)

after generalising to decade + pin area:
k-anonymity: 10   unique rows: 0

61.7% of rows are unique on three columns nobody would call personal data. k-anonymity is the size of the smallest group that shares the same quasi-identifier values; a table with k=1 offers no protection at all.

Generalising fixes the count and destroys resolution. Birth year became a decade, and the pin code became a broad area. That is the real trade, and it is the reason k-anonymity is unpopular for analytics workloads.

Common mistakes

Measuring the attack with average-case AUC only. Carlini et al. (2022) argue the meaningful question is the true-positive rate at very low false-positive rates — an attacker who confidently identifies 1% of the training set has done real damage while the AUC looks unremarkable. Report TPR at 0.1% FPR alongside AUC.

Assuming k-anonymity is enough. If all ten people in a group share the same diagnosis, the attacker learns it without identifying anybody. That gap is why l-diversity and t-closeness exist, and why neither fully closes it.

Sending training data to a third-party API for labelling or evaluation. That is a disclosure, whatever the terms of service say. Decide it deliberately, and write it into your model card.

Logging raw prompts forever. For an LLM product, the prompt log is usually a larger privacy liability than the model. Set a retention limit and redact before storage.

Believing embeddings are safe to share. Text embeddings can be partially inverted back to the source text. Treat an embedding as the data, not as a hash of it.

Try it yourself

Change min_samples_leaf to 5, 10, 20 and 40 and plot attack AUC against test accuracy. That curve is your privacy-utility trade-off, measured on your own model rather than assumed. Then add pin_code back into part two's generalised grouping and watch k collapse.

What to learn next

Researcher — Mathematics and papers.

Threat models

Precision about the adversary is the difference between a meaningful privacy claim and marketing.

  • Membership inference — decide whether record $z$ was in the training set $D$. Shokri et al. (2017) introduced the shadow-model attack; Yeom et al. (2018) showed a loss threshold alone is a strong baseline and formally tied advantage to the generalisation gap.
  • Attribute inference — recover a missing sensitive field of a partially known record.
  • Model inversion — reconstruct a representative input for a class. Fredrikson et al. (2015) recovered recognisable face images from a facial-recognition API.
  • Training-data extraction — recover verbatim sequences. Carlini et al. (2021) extracted memorised personal data from GPT-2; Carlini et al. (2023) showed extraction scales with model size and data duplication.
  • Property inference — infer a global property of the training distribution, such as the demographic mix, which can be commercially sensitive even when no individual is exposed.

The generalisation-gap bound

Yeom et al. (2018) formalise the relationship the developer block demonstrates. For a bounded loss and an attacker thresholding the loss, membership advantage is bounded by a function of the train-test loss gap. The practical consequence: a model with zero generalisation gap admits no loss-threshold membership attack, and every point of overfitting is a point of leakage.

This is a bound on one attack family, not on all attacks. It does not license "we do not overfit, therefore we are private".

Measuring correctly

Carlini et al. (2022), Membership Inference Attacks From First Principles, is the methodological correction the field needed. Two arguments:

  1. Average-case metrics (accuracy, AUC) obscure the attack that matters. An attack achieving 50.5% accuracy but a 1000× lift in TPR at 0.001 FPR is a serious breach reported as a non-event. Report full log-scale ROC curves.
  2. LiRA (Likelihood Ratio Attack) trains shadow models with and without the target record and performs a per-example hypothesis test on the loss distributions. It dominates prior attacks by orders of magnitude at low FPR, at the cost of training many shadow models.

Any privacy evaluation that reports only AUC against a global threshold should be treated as a lower bound on the true risk, and a loose one.

Memorisation in generative models

Feldman (2020), Does Learning Require Memorization?, gives a positive result that reframes the discussion: for long-tailed label distributions, memorising rare examples is necessary for near-optimal generalisation. Memorisation is not always a defect to be removed; it is sometimes load-bearing, which is why privacy costs utility.

Carlini et al. (2023), Quantifying Memorization Across Neural Language Models, establishes log-linear scaling of extractable memorisation with model capacity, example duplication count and prompt-context length. Deduplication of training data (Lee et al., 2022) is the single highest-leverage mitigation and also improves quality.

For diffusion models, Carlini et al. (2023), Extracting Training Data from Diffusion Models, recovered near-exact training images, with duplication again the dominant factor.

Defences, ordered by strength of guarantee

DefenceGuaranteeCost
Regularisation, early stoppingHeuristic; reduces one attack familyNear zero
Training-set deduplicationHeuristic; large empirical effect on extractionPreprocessing only
Confidence rounding, top-k API outputRaises attack cost onlyLow; defeats naive attacks
k-anonymity, l-diversitySyntactic, on the data releaseHigh utility loss
DP-SGDFormal $(\varepsilon, \delta)$ bound on any adversaryAccuracy, compute, tuning
Federated learning aloneNo formal guaranteeSystems complexity

The last row is the common misconception. Federated learning removes raw-data centralisation; gradients still leak. Zhu et al. (2019), Deep Leakage from Gradients, reconstruct training images from shared gradients. Federated learning composes with differential privacy and secure aggregation; on its own it is a data-locality property, not a privacy guarantee.

Papers

What to learn next