ROC vs precision-recall curves
ROC curves judge a ranking against all the negatives while precision-recall curves judge what the alarms contain — on rare positives the two tell very different stories.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A ROC curve asks "how well does the model separate the two groups overall?"
A precision-recall curve asks "when the model raises an alarm, how often is it right?" With rare positives, those are wildly different questions.
Think of an airport metal detector. One way to praise it: of all the harmless bags, how few made it beep? Sounds great — 99 in 100 harmless bags pass silently. The other way: of all the beeps today, how many were real threats? With one genuine threat per lakh of bags, nearly every beep is someone's belt buckle.
Same machine. Flattering answer from the first question, sobering answer from the second.
Why this matters
Most classifiers output a score, and you grade the ranking the scores produce. Two standard curves do the grading. The ROC curve plots how many true positives you catch against how many negatives you wrongly alarm, across every possible cut-off. The precision-recall curve plots, for the same cut-offs, how pure the alarms are against how many positives you caught.
The difference has teeth only when positives are rare — fraud, disease, defects. Then the mountain of negatives makes the ROC view generous. Wrongly alarming on 1% of a million negatives barely moves its false-alarm rate. It also buries your alert queue under ten thousand false alarms.
How it works
1,000,000 events, 100 truly positive
model alarms on 10,100 events:
catches 100 of 100 positives -> ROC view: brilliant
but 10,000 alarms are false -> precision: 100 / 10,100 = 1%
ROC saw: "only 1% of negatives alarmed"
PR saw: "99% of your alarms are junk"Neither curve is wrong. ROC answers a question about the model's general separating power, useful when both classes matter comparably. Precision-recall answers the operator's question — what will my alert queue contain? When classes are balanced, the two views largely agree. When positives are rare, believe the precision-recall view for any decision about deployment.
A real example you have seen
Email spam filtering runs at the other extreme — spam is common, both mistake types matter, and ROC-style thinking works fine. Now move the same maths to cancer screening, where one scan in a thousand is positive. Screening programmes report "how many recalls were real" — precision — because that number decides how much fear and cost each alarm spreads.
Remember this
- ROC: separation of the two groups overall; generous when negatives are plentiful.
- Precision-recall: what your alarms contain; harsh and honest on rare positives.
- The rarer the positive class, the more the PR view is the one that matters.
What to learn next
- Choosing a threshold from costs — collapsing the curve to the one point you will run at.
- Evaluating an outlier detector — these ideas under extreme imbalance.
- Model evaluation — the wider metric toolbox.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7.2 on CPU; last-digit drift elsewhere is normal.
One model, two report cards
A logistic regression on data with 1% positives:
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import average_precision_score, roc_auc_score
from sklearn.model_selection import train_test_split
X, y = make_classification(n_samples=40000, weights=[0.99], flip_y=0,
class_sep=1.0, n_informative=2, random_state=1)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.5, stratify=y, random_state=0)
p = LogisticRegression().fit(Xtr, ytr).predict_proba(Xte)[:, 1]
print(f"positives in the test set: {yte.mean():.1%}")
print("ROC-AUC:", round(roc_auc_score(yte, p), 3))
print("PR-AUC: ", round(average_precision_score(yte, p), 3))
print("PR-AUC of random guessing:", round(yte.mean(), 3))
top = np.argsort(p)[-100:] # the 100 most confident alarms
print("real positives among the top 100 alarms:", int(yte[top].sum()))positives in the test set: 1.0% ROC-AUC: 0.963 PR-AUC: 0.544 PR-AUC of random guessing: 0.01 real positives among the top 100 alarms: 74
The walkthrough
0.963 versus 0.544 — same model, same predictions. The ROC number invites celebration; the PR number counsels caution. The last line arbitrates: the 100 most confident alarms contain 74 real positives. Good, imperfect, and exactly what the PR view predicted you would feel in operation.
The baselines are the trap. Random guessing scores ROC-AUC 0.5 on any data, but PR-AUC equal to the positive rate — 0.01 here. So never compare a PR-AUC against 0.5, and never call 0.544 "barely better than a coin flip". It is 54 times the random baseline.
average_precision_score is the PR-AUC you want. It integrates precision over recall step-wise, without the optimistic interpolation some plotting tools apply. roc_auc_score has no such subtlety — one more reason ROC numbers travel between papers so smoothly.
stratify=y in the split keeps the 1% rate identical in both halves. With 400 positives total, an unlucky unstratified split visibly moves every number in this lesson.
Common mistakes
Reporting ROC-AUC alone on imbalanced problems. Reviewers and dashboards accept it; queues built on it drown analysts. Report PR-AUC alongside, and precision at your true alert budget — the detector-evaluation lesson makes this concrete for anomaly work.
Comparing PR-AUCs across datasets with different positive rates. The baseline moves with prevalence, so cross-dataset PR comparisons need the baseline stated next to each number.
Reading the curves but deploying a threshold nobody chose. Both curves sweep all thresholds; production runs at one. Choosing it belongs to cost-based reasoning, the next lesson.
Assuming a high ROC-AUC model can be fixed later. If the ranking is poor near the top — the region PR sees — no threshold rescues it. Ranking quality where you will operate is the property to select for.
Try it yourself
Change weights=[0.99] to weights=[0.5] for a balanced problem, rerun, and watch ROC-AUC and PR-AUC nearly agree. Then push to weights=[0.999] and watch them diverge further. Prevalence, not the model, controls the gap.
What to learn next
- Choosing a threshold from costs — collapsing the curve to the one point you will run at.
- Evaluating an outlier detector — these ideas under extreme imbalance.
- Model evaluation — the wider metric toolbox.
Researcher — Mathematics and papers.
Formal definitions
With score threshold $t$ swept over its range:
- ROC: $\big(\mathrm{FPR}(t), \mathrm{TPR}(t)\big)$, where $\mathrm{TPR} = \frac{TP}{TP+FN}$ (recall) and $\mathrm{FPR} = \frac{FP}{FP+TN}$.
- PR: $\big(\mathrm{Recall}(t), \mathrm{Precision}(t)\big)$, where $\mathrm{Precision} = \frac{TP}{TP+FP}$.
ROC-AUC equals $P(s^+ > s^-)$ for independent random positive and negative scores — the Mann–Whitney U statistic (Hanley and McNeil, 1982) — and is invariant to prevalence $\pi$ and to any monotone rescaling of scores. Precision, by contrast, depends on prevalence explicitly:
$$ \mathrm{Precision} = \frac{\pi \cdot \mathrm{TPR}}{\pi \cdot \mathrm{TPR} + (1 - \pi)\, \mathrm{FPR}} $$
Where $\pi$ is the positive prevalence. This identity is the whole lesson in one line: as $\pi \to 0$, precision is crushed by the $(1-\pi)\,\mathrm{FPR}$ term unless FPR shrinks proportionally — the regime where ROC's x-axis has almost no resolution left.
The correspondence theorem and its limits
Davis and Goadrich (2006), The relationship between precision-recall and ROC curves: for a fixed dataset, a curve dominates in ROC space iff it dominates in PR space. But AUC orderings can differ — model A can beat model B on ROC-AUC while losing on PR-AUC — because the two integrals weight threshold regions differently. ROC-AUC integrates uniformly over FPR; AP weights by where positives fall in the ranking, concentrating attention at the top.
Their second contribution: linear interpolation between PR points is invalid (precision is not linear in recall between operating points); correct interpolation goes through the implied confusion-matrix counts. average_precision_score uses the step-wise estimator precisely to avoid the optimistic trapezoid.
Estimation subtleties
- AP's estimator is biased upward at small positive counts; variance is dominated by $n_+$, so report bootstrap intervals over positives (Boyd et al., 2013 analyse PR confidence bands).
- ROC-AUC admits closed-form variance (DeLong et al., 1988), enabling proper paired tests between models — a genuine practical advantage of the ROC world.
- Partial AUCs restrict integration to relevant FPR (ROC) or recall (PR) ranges, answering "quality at the operating region" without abandoning curve summaries (McClish, 1989).
- Saito and Rehmsmeier (2015) demonstrate on biological data how ROC plots conceal what PR plots reveal at $\pi < 0.05$ — the standard citation when arguing this point with reviewers.
The decision-theoretic completion: neither AUC is a decision — both average over thresholds you will never use. Converting a ranking plus costs into one operating point is the subject of threshold choice, and proper scoring of the probabilities themselves belongs to calibration.
What to learn next
- Choosing a threshold from costs — collapsing the curve to the one point you will run at.
- Evaluating an outlier detector — these ideas under extreme imbalance.
- Model evaluation — the wider metric toolbox.