AI in Finance and Fraud

How a real fraud system is built

Fraud detection is an extremely rare-event problem solved in real time, where accuracy is close to meaningless and the right threshold depends on what a false alarm actually costs.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. The honest part
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A fraud system decides, in a fraction of a second, whether one transaction is the rare bad one.

Think about a ticket checker walking through a crowded train compartment. Almost every passenger has a valid ticket. Checking every single bag in detail would make the checker miss the next station. So they learn to move fast. They glance at a few things, and pull aside only the passengers who seem off. Sometimes they will stop an honest traveller by mistake. Sometimes they will miss a fare-dodger.

A card fraud system faces the identical trade-off, at thousands of transactions a second.

Why it exists

Fraud is genuinely rare. In most card portfolios, well under 1 in 100 transactions is fraudulent. The exact rate is guarded business information. It also varies enormously by merchant type and region. Treat any specific percentage you read online as illustrative, not universal. That rarity breaks the usual way of judging a model.

A system that predicts "not fraud" on every transaction would be right almost every time. It would also be completely useless. It would let every fraud through. This is why fraud detection cannot be judged by accuracy — how often the model is right overall. A model must instead be judged by how well it finds rare fraud, without drowning honest customers in false alarms.

There is also a hard time limit. A card payment must be approved or declined in around a second, at the point of sale. A fraud model that takes ten seconds to think is not a fraud model. It is a returned item and a frustrated customer.

How it works

transaction happens
        |
        v
  [ instant rule checks ]  -- e.g. "amount over the card's usual pattern"
        |
        v
  [ ML model scores the transaction, in milliseconds ]
        |
     +------+------+-------+
     |             |            |
 high-risk    borderline   high-confidence
   fraud         score       legitimate
     |             |            |
   BLOCK      human review     APPROVE
              queue, or a
              step-up check
              (OTP, extra
               verification)

Most production fraud systems are not "only a model." They are a stack. Instant hand-written rules catch known patterns. An ML model catches everything subtler. A human review queue handles the genuinely uncertain middle.

A real example you have seen

Your bank sometimes sends an SMS asking "was this you?" after an unusual purchase. That is the middle branch above — a step-up check. It gets used when the model is unsure. Blocking outright would upset too many honest customers, but letting it through silently is too risky.

The honest part

Fraud detection is a moving target. Every rule and every model gets learned by real fraudsters over time. They then adjust their behaviour to avoid it. A model that works well today needs to be retrained and re-checked regularly. This is not a "train once" problem. Any team that treats it that way will watch performance quietly decay.

Remember this

  • Fraud is rare, so a fraud model must be judged on precision and recall, never on plain accuracy.
  • The decision has to happen in near real time. A model too slow to run at checkout is not usable, however accurate.
  • Real fraud systems combine instant rules, a scoring model, and a human review queue — not a model alone.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Minimal runnable code

We build a synthetic transaction dataset with a realistic 2.5% fraud rate, and compare a lazy "always legitimate" baseline against a real model — using the right metrics, not accuracy.

fraud_imbalance_demo.py
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, precision_score, recall_score, average_precision_score

# Synthetic transactions: 10 numeric features, ~2.5% are fraud
X, y = make_classification(
    n_samples=5000, n_features=10, n_informative=5,
    weights=[0.98, 0.02], flip_y=0.01, random_state=0,
)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.3, stratify=y, random_state=0
)

print(f"fraud rate in test set: {y_test.mean():.3f}")
lazy_accuracy = 1 - y_test.mean()
print(f"lazy baseline (always predict 'legit') accuracy: {lazy_accuracy:.3f}")

model = RandomForestClassifier(n_estimators=200, class_weight="balanced", random_state=0)
model.fit(X_train, y_train)
pred = model.predict(X_test)
proba = model.predict_proba(X_test)[:, 1]

print(f"model accuracy:  {accuracy_score(y_test, pred):.3f}")
print(f"model precision: {precision_score(y_test, pred):.3f}")
print(f"model recall:    {recall_score(y_test, pred):.3f}")
print(f"average precision (PR-AUC): {average_precision_score(y_test, proba):.3f}")
Output
fraud rate in test set: 0.025
lazy baseline (always predict 'legit') accuracy: 0.975
model accuracy:  0.981
model precision: 0.765
model recall:    0.342
average precision (PR-AUC): 0.559

What actually happened

Read the accuracy numbers first: 97.5% for a model that catches zero fraud, versus 98.1% for a model that actually works. A 0.6-point accuracy gap is hiding a model that correctly flags 76.5% of the transactions it calls fraud (precision), and catches 34.2% of all real fraud (recall). Accuracy barely moved because legitimate transactions so heavily outnumber fraud that getting them all right dominates the score either way.

This is the single most important fact about fraud detection, made visible in two printed numbers.

  • weights=[0.98, 0.02] sets the synthetic fraud rate. Real fraud rates vary by merchant category and region — always ask your own data what the real rate is rather than assuming a number.
  • class_weight="balanced" tells the random forest to penalise mistakes on the rare class more heavily during training, a standard first move on imbalanced data — see class weights.
  • average_precision_score summarises the precision-recall trade-off across every possible threshold in one number, which is a fairer single-number summary here than ROC-AUC — see ROC vs precision-recall curves for why.

Common mistakes

Reporting accuracy as the headline number. On real fraud data with a 0.1-1% fraud rate, this mistake is even more severe than shown here — a lazy baseline can clear 99% accuracy while catching nothing.

Picking a single threshold without a cost model. A 0.5 probability cutoff has no special meaning. The right threshold depends on how much a missed fraud costs versus how much a false alarm annoys a customer — see choosing a threshold from costs.

Evaluating on a random split instead of a time-based one. Fraud patterns shift over time. A model tested on a random shuffle of historical data can look far better than it will on tomorrow's traffic, which is always in the future relative to training. Temporal leakage covers this in general.

Retraining rarely. A model frozen for a year is being tested, silently, by fraudsters who have a year to find its blind spots.

Try it yourself

Change weights=[0.98, 0.02] to weights=[0.995, 0.005], closer to some real-world card fraud rates. Re-run and watch the lazy baseline's accuracy climb even higher, while the actual work the model is doing (precision and recall) stays visible only in those two numbers.

What to learn next

Researcher — Mathematics and papers.

Why this is a cost-sensitive, not accuracy-sensitive, problem

Let c_FN be the cost of a missed fraud (the transaction amount, typically) and c_FP be the cost of a false alarm (customer friction, lost legitimate revenue, support cost). The optimal decision threshold t* on predicted probability p minimises expected cost, not error count:

t*  =  c_FP / (c_FP + c_FN)

Because c_FN (a stolen transaction, sometimes reversed at the bank's expense) is typically far larger than c_FP (one annoyed customer), the cost-optimal threshold sits well below 0.5 — the system should flag transactions the naive classifier would call "probably fine." This single asymmetry, not model architecture, drives most of the practical design of a fraud system.

Real-time constraints as a modelling constraint, not an afterthought

Card-present authorization typically has to complete within roughly one second end to end, of which the fraud model gets a small slice — often single-digit milliseconds once network and issuer processing are accounted for. This rules out models requiring expensive online feature computation (e.g. a full graph traversal per transaction) at decision time; such signals are instead precomputed asynchronously and looked up, not computed live. See latency and throughput for the general treatment of this kind of budget.

Delayed and noisy labels

A transaction's true label (fraud or not) is often only confirmed weeks later, via a customer dispute or a chargeback — and not every fraud is ever reported by the victim, so even "confirmed" labels undercount true fraud. This means:

  • Training labels available today reflect transactions from weeks ago, not yesterday — an unavoidable lag in a supervised setup.
  • Absence of a chargeback is evidence of "probably legitimate," not proof — a genuinely fraudulent transaction the cardholder never noticed is silently mislabelled as legitimate in most datasets.

Positive-unlabeled learning framings and conservative relabelling heuristics exist specifically to address this label noise; there is no fully clean solution, only mitigations.

Cost

Feature engineering for fraud typically dominates modelling effort: velocity features (transaction count/amount in the last N minutes for this card, merchant, or device), graph-derived features (shared device or address across accounts), and behavioural-biometric signals (typing rhythm, device fingerprint) each require dedicated, low-latency infrastructure to compute — feature stores exist largely to serve this exact need at production latency.

Key references

  • Dal Pozzolo, A., Boracchi, G., Caelen, O., Alippi, C., & Bontempi, G. (2018). Credit Card Fraud Detection: A Realistic Modeling and a Novel Learning Strategy. IEEE Transactions on Neural Networks and Learning Systems — addresses concept drift and delayed labels directly.
  • Bahnsen, A. C., Aouada, D., & Ottersten, B. (2015). Example-Dependent Cost-Sensitive Decision Trees. Expert Systems with Applications — formalises the cost-threshold relationship used above.

Current state

Most production fraud systems today combine a fast gradient-boosted or deep model for real-time scoring with a slower graph or sequence model run asynchronously to catch coordinated fraud rings, feeding back into the real-time model's features — a two-speed architecture directly shaped by the latency constraint. Purely synchronous "one model does everything" architectures are rare in serious deployments.

What to learn next