How a real fraud system is built
Fraud detection is an extremely rare-event problem solved in real time, where accuracy is close to meaningless and the right threshold depends on what a false alarm actually costs.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A fraud system decides, in a fraction of a second, whether one transaction is the rare bad one.
Think about a ticket checker walking through a crowded train compartment. Almost every passenger has a valid ticket. Checking every single bag in detail would make the checker miss the next station. So they learn to move fast. They glance at a few things, and pull aside only the passengers who seem off. Sometimes they will stop an honest traveller by mistake. Sometimes they will miss a fare-dodger.
A card fraud system faces the identical trade-off, at thousands of transactions a second.
Why it exists
Fraud is genuinely rare. In most card portfolios, well under 1 in 100 transactions is fraudulent. The exact rate is guarded business information. It also varies enormously by merchant type and region. Treat any specific percentage you read online as illustrative, not universal. That rarity breaks the usual way of judging a model.
A system that predicts "not fraud" on every transaction would be right almost every time. It would also be completely useless. It would let every fraud through. This is why fraud detection cannot be judged by accuracy — how often the model is right overall. A model must instead be judged by how well it finds rare fraud, without drowning honest customers in false alarms.
There is also a hard time limit. A card payment must be approved or declined in around a second, at the point of sale. A fraud model that takes ten seconds to think is not a fraud model. It is a returned item and a frustrated customer.
How it works
transaction happens
|
v
[ instant rule checks ] -- e.g. "amount over the card's usual pattern"
|
v
[ ML model scores the transaction, in milliseconds ]
|
+------+------+-------+
| | |
high-risk borderline high-confidence
fraud score legitimate
| | |
BLOCK human review APPROVE
queue, or a
step-up check
(OTP, extra
verification)Most production fraud systems are not "only a model." They are a stack. Instant hand-written rules catch known patterns. An ML model catches everything subtler. A human review queue handles the genuinely uncertain middle.
A real example you have seen
Your bank sometimes sends an SMS asking "was this you?" after an unusual purchase. That is the middle branch above — a step-up check. It gets used when the model is unsure. Blocking outright would upset too many honest customers, but letting it through silently is too risky.
The honest part
Fraud detection is a moving target. Every rule and every model gets learned by real fraudsters over time. They then adjust their behaviour to avoid it. A model that works well today needs to be retrained and re-checked regularly. This is not a "train once" problem. Any team that treats it that way will watch performance quietly decay.
Remember this
- Fraud is rare, so a fraud model must be judged on precision and recall, never on plain accuracy.
- The decision has to happen in near real time. A model too slow to run at checkout is not usable, however accurate.
- Real fraud systems combine instant rules, a scoring model, and a human review queue — not a model alone.
What to learn next
- Finding fraud rings with graphs — catching fraud that only shows up as a pattern between accounts.
- Class weights — one of the standard tools for training on rare-event data like this.
- ROC vs precision-recall curves — the right way to evaluate a model like this one.
Developer — Code and libraries.
Setup
pip install scikit-learnMinimal runnable code
We build a synthetic transaction dataset with a realistic 2.5% fraud rate, and compare a lazy "always legitimate" baseline against a real model — using the right metrics, not accuracy.
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, precision_score, recall_score, average_precision_score
# Synthetic transactions: 10 numeric features, ~2.5% are fraud
X, y = make_classification(
n_samples=5000, n_features=10, n_informative=5,
weights=[0.98, 0.02], flip_y=0.01, random_state=0,
)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3, stratify=y, random_state=0
)
print(f"fraud rate in test set: {y_test.mean():.3f}")
lazy_accuracy = 1 - y_test.mean()
print(f"lazy baseline (always predict 'legit') accuracy: {lazy_accuracy:.3f}")
model = RandomForestClassifier(n_estimators=200, class_weight="balanced", random_state=0)
model.fit(X_train, y_train)
pred = model.predict(X_test)
proba = model.predict_proba(X_test)[:, 1]
print(f"model accuracy: {accuracy_score(y_test, pred):.3f}")
print(f"model precision: {precision_score(y_test, pred):.3f}")
print(f"model recall: {recall_score(y_test, pred):.3f}")
print(f"average precision (PR-AUC): {average_precision_score(y_test, proba):.3f}")fraud rate in test set: 0.025 lazy baseline (always predict 'legit') accuracy: 0.975 model accuracy: 0.981 model precision: 0.765 model recall: 0.342 average precision (PR-AUC): 0.559
What actually happened
Read the accuracy numbers first: 97.5% for a model that catches zero fraud, versus 98.1% for a model that actually works. A 0.6-point accuracy gap is hiding a model that correctly flags 76.5% of the transactions it calls fraud (precision), and catches 34.2% of all real fraud (recall). Accuracy barely moved because legitimate transactions so heavily outnumber fraud that getting them all right dominates the score either way.
This is the single most important fact about fraud detection, made visible in two printed numbers.
weights=[0.98, 0.02]sets the synthetic fraud rate. Real fraud rates vary by merchant category and region — always ask your own data what the real rate is rather than assuming a number.class_weight="balanced"tells the random forest to penalise mistakes on the rare class more heavily during training, a standard first move on imbalanced data — see class weights.average_precision_scoresummarises the precision-recall trade-off across every possible threshold in one number, which is a fairer single-number summary here than ROC-AUC — see ROC vs precision-recall curves for why.
Common mistakes
Reporting accuracy as the headline number. On real fraud data with a 0.1-1% fraud rate, this mistake is even more severe than shown here — a lazy baseline can clear 99% accuracy while catching nothing.
Picking a single threshold without a cost model. A 0.5 probability cutoff has no special meaning. The right threshold depends on how much a missed fraud costs versus how much a false alarm annoys a customer — see choosing a threshold from costs.
Evaluating on a random split instead of a time-based one. Fraud patterns shift over time. A model tested on a random shuffle of historical data can look far better than it will on tomorrow's traffic, which is always in the future relative to training. Temporal leakage covers this in general.
Retraining rarely. A model frozen for a year is being tested, silently, by fraudsters who have a year to find its blind spots.
Try it yourself
Change weights=[0.98, 0.02] to weights=[0.995, 0.005], closer to some real-world card fraud rates. Re-run and watch the lazy baseline's accuracy climb even higher, while the actual work the model is doing (precision and recall) stays visible only in those two numbers.
What to learn next
- Class weights — the technique used above, explained in depth.
- Choosing a threshold from costs — turning a probability into an actual block/allow decision.
- Finding fraud rings with graphs — the next layer, for fraud that only shows up as a pattern across accounts.
Researcher — Mathematics and papers.
Why this is a cost-sensitive, not accuracy-sensitive, problem
Let c_FN be the cost of a missed fraud (the transaction amount, typically) and c_FP be the cost of a false alarm (customer friction, lost legitimate revenue, support cost). The optimal decision threshold t* on predicted probability p minimises expected cost, not error count:
t* = c_FP / (c_FP + c_FN)Because c_FN (a stolen transaction, sometimes reversed at the bank's expense) is typically far larger than c_FP (one annoyed customer), the cost-optimal threshold sits well below 0.5 — the system should flag transactions the naive classifier would call "probably fine." This single asymmetry, not model architecture, drives most of the practical design of a fraud system.
Real-time constraints as a modelling constraint, not an afterthought
Card-present authorization typically has to complete within roughly one second end to end, of which the fraud model gets a small slice — often single-digit milliseconds once network and issuer processing are accounted for. This rules out models requiring expensive online feature computation (e.g. a full graph traversal per transaction) at decision time; such signals are instead precomputed asynchronously and looked up, not computed live. See latency and throughput for the general treatment of this kind of budget.
Delayed and noisy labels
A transaction's true label (fraud or not) is often only confirmed weeks later, via a customer dispute or a chargeback — and not every fraud is ever reported by the victim, so even "confirmed" labels undercount true fraud. This means:
- Training labels available today reflect transactions from weeks ago, not yesterday — an unavoidable lag in a supervised setup.
- Absence of a chargeback is evidence of "probably legitimate," not proof — a genuinely fraudulent transaction the cardholder never noticed is silently mislabelled as legitimate in most datasets.
Positive-unlabeled learning framings and conservative relabelling heuristics exist specifically to address this label noise; there is no fully clean solution, only mitigations.
Cost
Feature engineering for fraud typically dominates modelling effort: velocity features (transaction count/amount in the last N minutes for this card, merchant, or device), graph-derived features (shared device or address across accounts), and behavioural-biometric signals (typing rhythm, device fingerprint) each require dedicated, low-latency infrastructure to compute — feature stores exist largely to serve this exact need at production latency.
Key references
- Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems — addresses concept drift and delayed labels directly.
- Bahnsen, A. C., Aouada, D., & Ottersten, B. (2015). Example-Dependent Cost-Sensitive Decision Trees. Expert Systems with Applications — formalises the cost-threshold relationship used above.
Current state
Most production fraud systems today combine a fast gradient-boosted or deep model for real-time scoring with a slower graph or sequence model run asynchronously to catch coordinated fraud rings, feeding back into the real-time model's features — a two-speed architecture directly shaped by the latency constraint. Purely synchronous "one model does everything" architectures are rare in serious deployments.
What to learn next
- Choosing a threshold from costs — the formal cost-threshold relationship.
- Feature stores — serving velocity and behavioural features at production latency.
- Finding fraud rings with graphs — the asynchronous, structural half of a real system.