AI in Finance and Fraud

Model risk management

Model risk management is the independent review, ongoing monitoring, and formal sign-off a financial model goes through before and after it touches real customers.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Model risk management checks a financial model before it is trusted with real money, and keeps checking it afterward.

Think about a restaurant health inspector. Before a restaurant opens, someone independent checks the kitchen against a standard. Not the chef, not the owner. The restaurant cannot legally open without passing. After it opens, inspections continue. A kitchen that was clean on opening day can slip later. The people cooking the food never get to declare the kitchen safe.

Every credit, fraud, and AML model in this section needs the same kind of independent check. It needs one before it goes near a real customer.

Why it exists

A model can decide whether to approve a loan, flag a fraud, or freeze an account. It can genuinely harm people if wrong: deny a fair loan, block a legitimate purchase, miss real financial crime. Model risk management, shortened to MRM, exists so these decisions are not left only to the model's own builders. However careful they are.

The core idea is independent validation. Someone who did not build the model checks it, using their own tests, before it is approved for use. This mirrors why credit decisions need explainable reasons. It mirrors why a backtest needs an untouched holdout. Never let the model builder be the model's only judge. Apply that same discipline at the level of an entire organisation.

Major financial regulators formally require this. Banks above a certain size must maintain a documented model inventory. They must tier each model by how much harm it could cause if wrong. Higher-risk models get deeper, more frequent independent review.

How it works

Model built
     |
     v
Documentation: what it does, what data it uses, its known limits
     |
     v
Independent validation team reviews it (NOT the builders)
     |
     v
Approved for a defined, limited use  ->  deployed
     |
     v
ONGOING monitoring: is it still behaving the way it did at approval?
     |
     v
Degraded performance, or the world changed  ->  re-review, retrain, or retire

Approval is never permanent. A model that was fine last year can quietly stop being fine as the world it operates in changes.

A real example you have seen

A bank sometimes pauses or "recalibrates" a credit or fraud model. This is often reported after unusual economic conditions, like a recession or a sudden change in fraud patterns. That pause is model risk management doing its job. It noticed the world changed. The model needs a fresh look before it is trusted further.

Remember this

  • A model should never be its own judge. Independent review, from people who did not build it, is the core idea.
  • Higher-stakes models (larger loans, bigger fraud impact) get deeper, more frequent scrutiny than lower-stakes ones.
  • Approval is not permanent. Ongoing monitoring exists because a model that was fine at launch can quietly stop being fine.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

One of the standard tools a model risk team uses to monitor a live model: the Population Stability Index (PSI), which measures how much a model's score distribution has shifted since launch.

psi_monitoring.py
import numpy as np

def population_stability_index(expected, actual, n_bins=10):
    quantiles = np.linspace(0, 1, n_bins + 1)
    breakpoints = np.quantile(expected, quantiles)
    breakpoints[0], breakpoints[-1] = -np.inf, np.inf  # catch any value outside the original range

    expected_counts, _ = np.histogram(expected, bins=breakpoints)
    actual_counts, _ = np.histogram(actual, bins=breakpoints)

    expected_pct = np.clip(expected_counts / len(expected), 1e-6, None)
    actual_pct = np.clip(actual_counts / len(actual), 1e-6, None)

    psi_per_bin = (actual_pct - expected_pct) * np.log(actual_pct / expected_pct)
    return psi_per_bin.sum()

rng = np.random.default_rng(0)
scores_at_launch = rng.normal(0.20, 0.08, 5000)     # model scores when the model launched
scores_now_stable = rng.normal(0.20, 0.08, 5000)    # today's scores, if nothing changed
scores_now_shifted = rng.normal(0.35, 0.10, 5000)   # today's scores, after real drift

psi_stable = population_stability_index(scores_at_launch, scores_now_stable)
psi_shifted = population_stability_index(scores_at_launch, scores_now_shifted)

print(f"PSI, stable population:  {psi_stable:.3f}  (below 0.1 = no real change)")
print(f"PSI, shifted population: {psi_shifted:.3f}  (above 0.25 = investigate before trusting the model)")
Output
PSI, stable population:  0.005  (below 0.1 = no real change)
PSI, shifted population: 2.242  (above 0.25 = investigate before trusting the model)

What actually happened

PSI compares two distributions of the same score, bucketed into the same bins, and measures how much probability mass moved between buckets. scores_now_stable was drawn from the exact same distribution as scores_at_launch, and PSI correctly reports almost no change — 0.005, far under the 0.1 "nothing to see here" line used across the credit and fraud industry.

scores_now_shifted has both a higher average (0.35 vs 0.20) and more spread (0.10 vs 0.08) than the launch distribution. PSI reports 2.242 — enormously above the 0.25 "investigate now" line. Something about the population being scored, or the world the model operates in, has genuinely changed.

  • breakpoints are built from the quantiles of expected (the reference distribution), not actual. This means each bin starts with roughly equal weight under expected, so a large PSI value can only come from actual no longer matching that shape.
  • Setting the first and last breakpoint to -inf/inf prevents new, more extreme values in actual from being silently dropped rather than counted.
  • np.clip(..., 1e-6, None) avoids log(0), which would otherwise crash the calculation the first time a bin happens to be empty.

Common mistakes

Monitoring only the final decision, not the underlying score. Two very different score distributions can produce the same approve/decline rate if a threshold happens to sit in a stable spot. PSI on the raw score catches shifts an approval-rate metric would miss entirely.

Picking bin edges from today's data instead of the reference period. This silently defeats the whole point — bins built from actual will always look roughly balanced against actual, hiding real drift. Bins must come from expected, fixed at approval time.

Treating a PSI alert as proof the model is wrong. A high PSI means the world looks different than it did at approval. That can mean the model needs retraining, or it can mean a real, expected seasonal shift. Either way, it is a signal to investigate, not an automatic verdict.

Only checking PSI once, at deployment. The entire value of this metric is in running it on a schedule, against fresh live data, for as long as the model is in use.

Try it yourself

Change scores_now_shifted to rng.normal(0.22, 0.08, 5000) — a small, realistic shift rather than a dramatic one. Watch where the PSI value lands relative to the 0.1 and 0.25 lines, and notice how much smaller a real, gradual shift looks compared to the dramatic example above.

What to learn next

Researcher — Mathematics and papers.

PSI, formally, and its limits

PSI  =  SUM_i  ( A_i - E_i ) * ln( A_i / E_i )
  • E_i — proportion of the reference (expected) population in bin i
  • A_i — proportion of the current (actual) population in bin i

This is the symmetrised form of the KL divergence between the two binned distributions, and shares its interpretation as an information-theoretic measure of how much one distribution has diverged from the other. It is not a formal statistical test with a p-value; the conventional 0.1/0.25 thresholds are industry heuristics, not derived from a null distribution, and their appropriateness depends on bin count and sample size, neither of which the raw PSI number reflects.

PSI also only measures marginal distribution shift on the scored variable — it says nothing about whether the model's relationship between inputs and true outcomes has changed (concept drift), only whether the population being scored looks different (covariate shift). A model's discrimination and calibration can degrade even with a stable PSI, if the underlying relationship between features and default has shifted while the feature distribution has not.

The three-lines-of-defence model

Most large financial institutions structure model governance around three formally separated functions:

  1. First line — the team that builds and owns the model, responsible for its day-to-day performance.
  2. Second line — independent model validation and risk management, responsible for challenging and approving the model, structurally separated from the first line's reporting chain.
  3. Third line — internal audit, responsible for checking that the first and second lines are actually following the documented process, not only that a model exists.

This structure, formalised for large US banks under Federal Reserve/OCC SR 11-7 (2011) and echoed in the EU's guidelines and the UK PRA's SS1/23, exists specifically to prevent the conflict of interest in a model builder validating their own work.

Model risk tiering

Institutions typically classify each model into a risk tier based on materiality (financial exposure if wrong) and complexity (how hard the model is to fully understand and test), with tier determining validation depth and revalidation frequency. A small internal model used for reporting might be revalidated every few years; a primary credit decisioning model touching billions in exposure is typically revalidated annually or on any material change, with continuous monitoring in between.

Cost

Formal model risk management is a substantial, ongoing operational cost — dedicated independent validation staff, documentation standards, monitoring infrastructure, and a full model inventory system — which is precisely why it is concentrated at regulated financial institutions handling the highest-stakes decisions, rather than applied uniformly across every ML system a company runs. The cost is deliberate: it is priced against the cost of an undetected model failure at scale.

Key references

  • Board of Governors of the Federal Reserve System & OCC (2011). SR 11-7: Supervisory Guidance on Model Risk Management.
  • UK Prudential Regulation Authority (2023). SS1/23: Model Risk Management Principles for Banks.
  • European Central Bank (2024). Guide on Effective Risk Data Aggregation and Risk Reporting and related ECB model risk guidance for significant institutions.
  • Siddiqi, N. (2017). Intelligent Credit Scoring (2nd ed.). Wiley — chapter on ongoing monitoring, including PSI in practical use.

Current state

Regulatory guidance written before large-scale deployment of gradient-boosted and deep models is being actively reinterpreted for these harder-to-validate model classes; explainability tooling (SHAP, surrogate models) is now commonly required as part of the validation package for any model whose internals are not directly interpretable, precisely to satisfy an MRM framework that predates them.

What to learn next