Reject inference
A credit model only ever learns from applicants who were approved, so it never sees what would have happened with the people it would have rejected — and that blind spot has a name.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Reject inference means guessing what would have happened to loan applicants who were never approved. Nobody ever really finds out.
Think about that kirana shop owner who lends groceries on credit. He refuses a stranger who looks unreliable. He never learns whether that stranger would have paid every rupee back. The stranger never got the chance. Every lesson he learns about "who repays" comes only from people he already trusted.
A bank's historical data has exactly the same hole in it, at a much larger scale.
Why it exists
A new credit model trains on a bank's own history: who applied, who was approved, and who repaid or defaulted. But repayment outcomes only exist for people who were approved. Applicants rejected in the past have no outcome recorded at all. They never got the loan, so there is nothing to observe.
This creates a quiet trap. A model trained only on approved applicants learns "who repays, among people we already chose to trust." It has never seen an outcome for the kind of applicant who used to get rejected. Suppose the bank now wants to loosen its policy. It would be asking the model to judge people it has zero real evidence about.
Worse, old approval decisions were often not based on income and history alone. A loan officer's gut feeling may have quietly influenced who got approved. So might a personal relationship, or an unrecorded reference check. These are factors the new model never sees as a feature, because nobody wrote them down.
How it works
Every past applicant
|
v
Historical approval decision (income, but also officer judgement)
|
+---------+----------+
| |
Approved Rejected
| |
Loan given No loan given
| |
Outcome KNOWN Outcome NEVER KNOWN
(repaid / defaulted) (nothing to learn from)A model trained only on the left branch is trained on a selected sample. It is not a random slice of "everyone who applied." It is a slice shaped by decisions the model never sees.
A real example you have seen
Banks want to expand lending to people with thin or no credit history. Their existing model was trained on people who already had enough history to get approved before. There is no honest way to know how the "no history" group would perform. The only real way to find out is to lend cautiously to a small, monitored batch of them. Then watch what actually happens.
The honest part
There is no trick that fully solves reject inference. Every technique is a set of assumptions about the invisible group, not a way to actually observe them. Some methods give rejected applicants estimated labels. Others reweight the approved sample, to look more like the full applicant pool. Both reduce the bias. None of them remove it. Any credit model expanding into previously-rejected territory needs conservative pilots and human review. It should never get blind trust in the number it prints.
Remember this
- A credit model only ever learns from approved applicants. Rejected applicants leave no outcome to learn from.
- This blind spot can make a model overconfident about people who resemble the rejected group. It has no real evidence either way.
- Expanding a credit policy safely means small, monitored pilots. Never a single leap based on untested extrapolation.
What to learn next
- Explaining a credit decision to a regulator — what a lender must do once a decision, approve or deny, is made.
- Group leakage — a related way training data silently fails to represent reality.
- Fairness metrics — checking whether this blind spot falls unevenly across groups of people.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyMinimal runnable code
This is a synthetic simulation. It plays god for a moment — generating true outcomes for everyone, including people who would have been rejected — purely so we can show the size of the blind spot. A real bank can never see this second, "true" column; the point of this code is to show why that matters.
import numpy as np
from sklearn.linear_model import LogisticRegression
rng = np.random.default_rng(3)
n = 4000
income = rng.normal(50, 15, n) # income in thousands -- the only feature our model gets
stability = rng.normal(0, 1, n) # a factor loan officers judged by eye, never digitised
# True default risk depends on BOTH income and the hidden stability factor
logit = 0.08 * (40 - income) - 1.2 * stability
true_risk = 1 / (1 + np.exp(-logit))
defaulted = rng.binomial(1, true_risk)
# Historical approvals: officers let in some low-income applicants who SEEMED stable
approved = (income > 45) | ((income > 30) & (stability > 1.0))
print("historical approval rate:", round(approved.mean(), 2))
# Our model only ever has income to work with -- stability was never recorded
X_obs = income[approved].reshape(-1, 1)
y_obs = defaulted[approved]
model_observed = LogisticRegression().fit(X_obs, y_obs)
# The truth a bank normally CANNOT see: real outcomes at every income level
model_full = LogisticRegression().fit(income.reshape(-1, 1), defaulted)
test_income = np.array([[30], [35], [40], [50]])
print()
print("income trained-on-approved-only true population rate")
for inc, po, pf in zip(test_income.ravel(),
model_observed.predict_proba(test_income)[:, 1],
model_full.predict_proba(test_income)[:, 1]):
print(f"{inc:>6} {po:.2f} {pf:.2f}")historical approval rate: 0.68
income trained-on-approved-only true population rate
30 0.48 0.64
35 0.43 0.57
40 0.38 0.49
50 0.29 0.34What actually happened
At every income level, the model trained only on approved applicants underestimates true default risk — and the gap is worst exactly where it matters most, at low income, where the bank is most likely to be considering an expansion.
The reason is stability. It genuinely affects default risk, and it quietly influenced who got approved historically. Among approved low-income applicants, stability is disproportionately high — the low-income people who got in were the ones who seemed safe by some unrecorded judgement. Because stability was never turned into a feature, the observed-only model cannot see that it is being shown a cherry-picked slice of low-income applicants, not a representative one.
model_fullexists only to make this lesson's point. No bank has this in real life — it requires knowing outcomes for people who were never given a loan.- The gap between the two models is not random noise. It has the same sign and grows in the same direction (worse underestimation at lower income) every time this script runs with this seed, because it is a structural bias, not a fluke.
Common mistakes
Assuming a bigger dataset fixes this. More approved applicants does not add a single rejected applicant's true outcome. The blind spot does not shrink with more data of the same kind.
Treating a reject-inference technique as ground truth. Methods like assigning "soft" inferred labels to rejected applicants (fuzzy augmentation) or reweighting the approved sample are documented, standard techniques — and every one of them is still a guess dressed up as a number. Track their assumptions, don't trust their output blindly.
Expanding approval policy in one big step. Because the model's extrapolation into "never approved before" territory is unverified, a real deployment tests it on a small, monitored batch first, and only expands once real outcomes confirm the extrapolation was reasonable.
Forgetting this applies beyond credit. Any model trained on "people who were selected by an old process" — hiring models trained on past hires, medical models trained on patients who were tested — inherits the same blind spot.
Try it yourself
Change stability > 1.0 to stability > 0.3 in the approval rule, which loosens the historical policy and approves a less selectively "stable" group of low-income applicants. Re-run the script and watch the gap between the two models shrink — because the approved sample is now closer to representative of the true population at that income level.
What to learn next
- Fairness metrics — checking whether this bias lands unevenly on particular groups.
- Group leakage — the general pattern of a model learning from a slice that quietly does not represent the whole.
- Explaining a credit decision to a regulator — the next step in the credit pipeline.
Researcher — Mathematics and papers.
Formal framing as sample selection
Let Y be the true default outcome and A be the historical approval indicator. Standard supervised learning assumes the training sample is drawn from P(Y | X). Credit data is instead drawn from P(Y | X, A = 1).
By Bayes' rule:
P(Y | X, A=1) = P(A=1 | X, Y) * P(Y | X) / P(A=1 | X)P(A=1 | X, Y)— the probability of historical approval given features and the (usually unknown at approval time) eventual outcome- If approval depended only on
X(missing at random, MAR), this ratio does not depend onY, andP(Y | X, A=1) = P(Y | X)— no bias, reject inference is unnecessary. - If approval also depended on an unobserved factor correlated with
Y(missing not at random, MNAR, the case simulated above viastability), the equality breaks, and every model trained onA=1data is estimating a biased quantity.
This is structurally identical to the Heckman (1979) sample-selection problem in econometrics, and reject inference techniques are largely re-derivations of solutions built for that literature.
Standard reject-inference techniques
Hard cutoff augmentation. Assign rejected applicants a label using the current model's own prediction, then retrain including them. This provably cannot correct MNAR bias — it re-teaches the model its own belief, which can amplify existing bias under repeated iteration (Anderson, 2007, calls this reinforcement rather than correction).
Fuzzy augmentation. Assign each rejected applicant a fractional "soft" weight split between good/bad outcomes, proportional to the current model's predicted probability, and retrain on the augmented, larger sample. Reduces variance of the estimate but shares the same MNAR limitation as hard augmentation.
Reweighting (inverse probability weighting). Model P(A=1 | X) — the propensity to be approved — separately, then upweight approved observations that resemble commonly-rejected profiles when fitting the outcome model. This corrects for selection on X but not on unobserved factors correlated with Y.
Bureau-score augmentation. Where legally and contractually permitted, obtain outcome data on rejected applicants from a credit bureau (they may have gone on to borrow elsewhere). This is the only listed technique that uses real, not inferred, outcomes — and it is often unavailable or heavily restricted.
Why none of these are a full solution
Every documented reject-inference method assumes something about the unobserved mechanism behind historical rejections, and that assumption is unverifiable using the biased sample itself. Crook & Banasik (2004) show formally that the bias from MNAR selection can be bounded but not eliminated without external data — the "true population" model in this lesson's code exists precisely because it is otherwise unobservable.
Cost
The only method that resolves the bias without unverifiable assumptions is a randomised champion-challenger or exploration policy: approving a small, randomised fraction of otherwise-rejected applicants, at a controlled and monitored financial cost, purely to collect ground truth. This is standard in mature credit operations and is the credit-industry version of exploration in a bandit problem.
Key references
- Heckman, J. J. (1979). Sample Selection Bias as a Specification Error. Econometrica 47(1).
- Anderson, R. (2007). The Credit Scoring Toolkit. Oxford University Press — chapter 18 covers augmentation techniques and their limits in practical detail.
- Crook, J., & Banasik, J. (2004). Does Reject Inference Really Improve the Performance of Application Scoring Models? Journal of Banking & Finance.
Current state
There is no consensus best method; industry practice mixes reweighting with small-scale randomised approval experiments where regulation and risk appetite allow it. Any paper or vendor claiming a reject-inference technique "solves" the selection bias should be read against Crook & Banasik's finding that measured performance gains from these techniques are frequently smaller than practitioners expect, and sometimes negative.
What to learn next
- Fairness metrics — quantifying whether this selection bias disadvantages specific groups.
- Colliders and selection bias — the causal-inference framing of exactly this problem.
- Model risk management — where an unverified reject-inference assumption must be documented and reviewed.