Privacy in machine learning
A trained model can leak the people it was trained on — and both the leak and the fix are measurable in a few lines of code.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A model remembers some of the people it learned from, and a stranger can sometimes tell who they were.
The analogy you have already lived
Think of a friend who read your diary once, years ago. Ask them about your life and they mostly speak in vague generalities.
But say a very specific sentence from that diary and watch their face. The flicker of recognition tells you they have seen it before. They never quoted a line. They gave it away by reacting differently to something familiar.
A model does the same flicker. It answers with more confidence about the exact examples it was trained on. That confidence gap is a leak, and it can be measured.
Why this is a real problem, not a theoretical one
Deleting names does not make data anonymous. This is the single most expensive misunderstanding in the field.
Suppose a hospital publishes patient records with the names removed. Each row still has a birth year, a pin code and a gender. In a city of millions, that combination alone identifies most people uniquely.
Anyone holding a voter roll can match the rows back to names. The diagnosis was the only thing hidden, and now it is not.
The developer block measures this on a fake table of 5,000 people. More than 60% of the rows are unique on those three harmless-looking columns.
The three ways a model leaks
It tells you who was in the training data. By answering more confidently on training examples. This is called membership inference — working out whether a specific person's record was used.
It repeats things word for word. A model trained on emails can be prompted into reproducing an address or a phone number it saw during training.
It lets you reconstruct a face or a record. From enough queries, an attacker can rebuild an approximation of an individual training example.
The picture
real people → training data → model → answers
▲ │
└────────── attacker works backwards ────┘
"Was Ramesh's record used to train this?"
"What phone number follows this name?"The arrow going backwards is the whole problem. Most teams design only the forward arrow.
The good news, which is genuinely good
The leak is caused by memorising. A model that memorises its training set is both worse at its job and leakier.
So the same discipline that stops overfitting — a model memorising its training examples instead of learning the pattern — also reduces the leak. In the experiment below, one settings change drops the attack's success from 93% to 55%, and costs one percentage point of accuracy. That is an unusually cheap trade.
What is honestly hard here
There is no switch labelled "private". Every real defence costs some accuracy, and the amount is not knowable in advance.
And privacy is not only a model problem. Logs, backups, prompts sent to a third-party API, and screenshots in a support ticket all leak. The model is often the safest part of the system.
Remember this
- Deleting names does not anonymise data — combinations of ordinary columns identify people.
- Models leak by being more confident about their training examples.
- Fighting memorisation is the same work as fighting overfitting, and it is cheap.
What to learn next
- Differential privacy — the one defence that comes with a proof.
- Federated learning — training without collecting the data centrally.
- Overfitting and underfitting — the same failure, seen from the accuracy side.
Developer — Code and libraries.
Measure the leak before you argue about it
Privacy in ML becomes an engineering problem the moment you can put a number on it. The number is membership inference attack AUC: how well can an attacker tell members of the training set from non-members, using only model outputs?
An AUC of 0.5 means the attacker learns nothing. An AUC of 0.9 means they are nearly always right.
Setup
pip install numpy scikit-learn pandasRuns on CPU in a couple of seconds.
Part 1 — a membership inference attack in fifteen lines
import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(21)
n, d = 1000, 20
X = rng.normal(0, 1, (n, d))
y = (X[:, :3].sum(1) + rng.normal(0, 2.0, n) > 0).astype(int) # deliberately noisy labels
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, random_state=0)
def attack(model, label):
model.fit(Xtr, ytr)
# the attacker only needs the probability the model gives to the TRUE label
p_member = model.predict_proba(Xtr)[np.arange(len(ytr)), ytr]
p_outsider = model.predict_proba(Xte)[np.arange(len(yte)), yte]
scores = np.r_[p_member, p_outsider]
is_member = np.r_[np.ones(len(ytr)), np.zeros(len(yte))]
print(f"{label:26s} train_acc={model.score(Xtr, ytr):.3f} "
f"test_acc={model.score(Xte, yte):.3f} "
f"mean_conf_member={p_member.mean():.3f} "
f"mean_conf_outsider={p_outsider.mean():.3f} "
f"ATTACK_AUC={roc_auc_score(is_member, scores):.3f}")
print("An attack AUC of 0.500 means the attacker learns nothing.\n")
attack(RandomForestClassifier(n_estimators=200, random_state=0), "memorising model")
attack(RandomForestClassifier(n_estimators=200, min_samples_leaf=40, random_state=0),
"regularised model")An attack AUC of 0.500 means the attacker learns nothing. memorising model train_acc=1.000 test_acc=0.744 mean_conf_member=0.852 mean_conf_outsider=0.603 ATTACK_AUC=0.932 regularised model train_acc=0.784 test_acc=0.734 mean_conf_member=0.582 mean_conf_outsider=0.562 ATTACK_AUC=0.554
What that result actually says
The first model gives away its training set. Attack AUC 0.932. Note the two confidence columns: 0.852 for members against 0.603 for outsiders. The model is not quoting anybody's data. It is reacting differently, and that is enough.
The gap between train and test accuracy is the leak. 1.000 against 0.744 is a 25-point generalisation gap, and the attack AUC tracks it closely. Overfitting and privacy leakage are the same phenomenon viewed from two sides.
One hyperparameter closed most of the hole. min_samples_leaf=40 forces every leaf to describe at least 40 people, so no leaf is about one person. Attack AUC falls to 0.554. Test accuracy falls from 0.744 to 0.734.
But 0.554 is not 0.500. Regularisation reduces the leak; it does not bound it. For a bound you need differential privacy, which is the next lesson.
Part 2 — "anonymised" data, measured
import numpy as np, pandas as pd
rng = np.random.default_rng(2)
N = 5000
df = pd.DataFrame({
"name_removed": ["<redacted>"] * N, # the "anonymisation" people trust
"birth_year": rng.integers(1955, 2006, N),
"pin_code": rng.integers(400001, 400105, N), # ~104 Mumbai pin codes
"gender": rng.choice(["F", "M"], N),
"diagnosis": rng.choice(["asthma", "diabetes", "none"], N),
})
quasi = ["birth_year", "pin_code", "gender"]
sizes = df.groupby(quasi).size()
print(f"rows: {N} quasi-identifiers: {quasi}")
print(f"k-anonymity of this table (smallest group): {sizes.min()}")
print(f"rows that are UNIQUE on those three columns: {(sizes == 1).sum()} "
f"({(sizes == 1).sum() / N:.1%})")
# Generalise: decade instead of year, first 4 digits of the pin code
df["decade"] = (df.birth_year // 10) * 10
df["pin_area"] = df.pin_code // 100
sizes2 = df.groupby(["decade", "pin_area", "gender"]).size()
print(f"\nafter generalising to decade + pin area:")
print(f"k-anonymity: {sizes2.min()} unique rows: {(sizes2 == 1).sum()}")rows: 5000 quasi-identifiers: ['birth_year', 'pin_code', 'gender'] k-anonymity of this table (smallest group): 1 rows that are UNIQUE on those three columns: 3087 (61.7%) after generalising to decade + pin area: k-anonymity: 10 unique rows: 0
61.7% of rows are unique on three columns nobody would call personal data. k-anonymity is the size of the smallest group that shares the same quasi-identifier values; a table with k=1 offers no protection at all.
Generalising fixes the count and destroys resolution. Birth year became a decade, and the pin code became a broad area. That is the real trade, and it is the reason k-anonymity is unpopular for analytics workloads.
Common mistakes
Measuring the attack with average-case AUC only. Carlini et al. (2022) argue the meaningful question is the true-positive rate at very low false-positive rates — an attacker who confidently identifies 1% of the training set has done real damage while the AUC looks unremarkable. Report TPR at 0.1% FPR alongside AUC.
Assuming k-anonymity is enough. If all ten people in a group share the same diagnosis, the attacker learns it without identifying anybody. That gap is why l-diversity and t-closeness exist, and why neither fully closes it.
Sending training data to a third-party API for labelling or evaluation. That is a disclosure, whatever the terms of service say. Decide it deliberately, and write it into your model card.
Logging raw prompts forever. For an LLM product, the prompt log is usually a larger privacy liability than the model. Set a retention limit and redact before storage.
Believing embeddings are safe to share. Text embeddings can be partially inverted back to the source text. Treat an embedding as the data, not as a hash of it.
Try it yourself
Change min_samples_leaf to 5, 10, 20 and 40 and plot attack AUC against test accuracy. That curve is your privacy-utility trade-off, measured on your own model rather than assumed. Then add pin_code back into part two's generalised grouping and watch k collapse.
What to learn next
- Differential privacy — the one defence that comes with a proof.
- Federated learning — training without collecting the data centrally.
- Overfitting and underfitting — the same failure, seen from the accuracy side.
Researcher — Mathematics and papers.
Threat models
Precision about the adversary is the difference between a meaningful privacy claim and marketing.
- Membership inference — decide whether record $z$ was in the training set $D$. Shokri et al. (2017) introduced the shadow-model attack; Yeom et al. (2018) showed a loss threshold alone is a strong baseline and formally tied advantage to the generalisation gap.
- Attribute inference — recover a missing sensitive field of a partially known record.
- Model inversion — reconstruct a representative input for a class. Fredrikson et al. (2015) recovered recognisable face images from a facial-recognition API.
- Training-data extraction — recover verbatim sequences. Carlini et al. (2021) extracted memorised personal data from GPT-2; Carlini et al. (2023) showed extraction scales with model size and data duplication.
- Property inference — infer a global property of the training distribution, such as the demographic mix, which can be commercially sensitive even when no individual is exposed.
The generalisation-gap bound
Yeom et al. (2018) formalise the relationship the developer block demonstrates. For a bounded loss and an attacker thresholding the loss, membership advantage is bounded by a function of the train-test loss gap. The practical consequence: a model with zero generalisation gap admits no loss-threshold membership attack, and every point of overfitting is a point of leakage.
This is a bound on one attack family, not on all attacks. It does not license "we do not overfit, therefore we are private".
Measuring correctly
Carlini et al. (2022), Membership Inference Attacks From First Principles, is the methodological correction the field needed. Two arguments:
- Average-case metrics (accuracy, AUC) obscure the attack that matters. An attack achieving 50.5% accuracy but a 1000× lift in TPR at 0.001 FPR is a serious breach reported as a non-event. Report full log-scale ROC curves.
- LiRA (Likelihood Ratio Attack) trains shadow models with and without the target record and performs a per-example hypothesis test on the loss distributions. It dominates prior attacks by orders of magnitude at low FPR, at the cost of training many shadow models.
Any privacy evaluation that reports only AUC against a global threshold should be treated as a lower bound on the true risk, and a loose one.
Memorisation in generative models
Feldman (2020), Does Learning Require Memorization?, gives a positive result that reframes the discussion: for long-tailed label distributions, memorising rare examples is necessary for near-optimal generalisation. Memorisation is not always a defect to be removed; it is sometimes load-bearing, which is why privacy costs utility.
Carlini et al. (2023), Quantifying Memorization Across Neural Language Models, establishes log-linear scaling of extractable memorisation with model capacity, example duplication count and prompt-context length. Deduplication of training data (Lee et al., 2022) is the single highest-leverage mitigation and also improves quality.
For diffusion models, Carlini et al. (2023), Extracting Training Data from Diffusion Models, recovered near-exact training images, with duplication again the dominant factor.
Defences, ordered by strength of guarantee
| Defence | Guarantee | Cost |
|---|---|---|
| Regularisation, early stopping | Heuristic; reduces one attack family | Near zero |
| Training-set deduplication | Heuristic; large empirical effect on extraction | Preprocessing only |
| Confidence rounding, top-k API output | Raises attack cost only | Low; defeats naive attacks |
| k-anonymity, l-diversity | Syntactic, on the data release | High utility loss |
| DP-SGD | Formal $(\varepsilon, \delta)$ bound on any adversary | Accuracy, compute, tuning |
| Federated learning alone | No formal guarantee | Systems complexity |
The last row is the common misconception. Federated learning removes raw-data centralisation; gradients still leak. Zhu et al. (2019), Deep Leakage from Gradients, reconstruct training images from shared gradients. Federated learning composes with differential privacy and secure aggregation; on its own it is a data-locality property, not a privacy guarantee.
Papers
- Shokri et al., Membership Inference Attacks Against Machine Learning Models, S&P 2017 — arxiv.org/abs/1610.05820
- Yeom et al., Privacy Risk in Machine Learning, CSF 2018 — arxiv.org/abs/1709.01604
- Carlini et al., Membership Inference Attacks From First Principles, S&P 2022 — arxiv.org/abs/2112.03570
- Carlini et al., Extracting Training Data from Large Language Models, USENIX Security 2021 — arxiv.org/abs/2012.07805
- Feldman, Does Learning Require Memorization?, STOC 2020 — arxiv.org/abs/1906.05271
- Zhu, Liu and Han, Deep Leakage from Gradients, NeurIPS 2019 — arxiv.org/abs/1906.08935
- Sweeney, k-Anonymity: A Model for Protecting Privacy, 2002
What to learn next
- Differential privacy — the one defence that comes with a proof.
- Federated learning — training without collecting the data centrally.
- Overfitting and underfitting — the same failure, seen from the accuracy side.