Explaining a credit decision to a regulator
A lender must give a declined applicant specific, checkable reasons, not only a probability number — which means the model's output has to be turned into words a human can verify.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
When a loan is declined, the lender must be able to say specifically why — not only show a number.
Picture that kirana shop owner one more time. A regular customer he usually trusts asks for credit. Today he says no. "A feeling" is not good enough for a long-standing customer who will ask why. He has to say something checkable. "You were late twice last month, and you already owe me for two weeks." A real reason, not a shrug.
Banks are held to the same standard, by law, at far greater scale.
Why it exists
A model that says "probability of default: 0.31, decline" tells the applicant nothing they can act on. It cannot be checked, argued with, or learned from. Regulators in most countries require something better. They require specific, individual reasons for a declined application, often called adverse action reasons or reason codes.
This exists for two reasons. First, an applicant deserves to know what to fix. Pay down debt, wait for a longer credit history, correct a bureau error. Second, specific reasons make it possible to check something after the fact. Is a model quietly discriminating against a protected group? Vague reasons are impossible to audit.
This is where model choice starts to matter as much as model accuracy. A model whose decision breaks cleanly into per-feature reasons is usable here. A model that cannot explain itself may be unusable for this exact reason, however accurate it is. This is a regulated environment. Accuracy alone is not the whole job.
How it works
applicant declined
|
v
model's internal reasoning, broken into per-feature pieces
|
v
ranked list: which factors pushed the decision most
|
v
"Your application was declined mainly because of:
1. Outstanding debt relative to income
2. Missed payments in the past 12 months"Each reason must correspond to something real and checkable in the applicant's data. It cannot be a vague label like "risk factor 7."
A real example you have seen
A rejected credit card or loan application in India comes with a reason. It often references your CIBIL report — a low score, a missed EMI, a high credit utilisation ratio. That written reason is the direct product of this requirement. The underlying model might be a simple scorecard, or something more complex behind the scenes.
Remember this
- A declined applicant has a right, in most jurisdictions, to specific and checkable reasons — not only a number.
- This is a real constraint on model choice. A model that cannot produce individual reasons may not be usable for the final decision, however accurate it is.
- The reasons exist to protect the applicant and to make discrimination auditable, not as a courtesy.
What to learn next
- How a real fraud system is built — a different finance model, and a different explainability standard.
- SHAP and LIME — the general-purpose tools this lesson's code is a hand-built miniature of.
- Fairness metrics — using consistent reasons to check for uneven treatment across groups.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyMinimal runnable code
For a logistic regression model, "why did this applicant score this way" has an exact, computable answer, because the model's output is a sum of per-feature pieces.
import numpy as np
from sklearn.linear_model import LogisticRegression
feature_names = ["income", "debt", "missed_payments"]
X = np.array([
[80, 5, 0], [65, 10, 1], [120, 0, 0], [30, 20, 3], [45, 15, 2],
[90, 8, 0], [25, 25, 4], [110, 2, 0], [55, 12, 1], [35, 18, 3],
[70, 6, 0], [40, 22, 3], [100, 4, 0], [50, 14, 2], [60, 9, 1],
[28, 24, 4], [95, 3, 0], [38, 19, 2], [75, 7, 0], [48, 16, 2],
])
y = np.array([0,1,0,1,1, 0,1,0,0,1, 0,1,0,1,0, 1,0,1,0,1])
model = LogisticRegression()
model.fit(X, y)
mean_applicant = X.mean(axis=0)
applicant = np.array([32, 21, 3]) # a declined case
# Each feature's push away from a TYPICAL applicant's log-odds,
# in the exact units the model itself reasons in.
contributions = model.coef_[0] * (applicant - mean_applicant)
print(f"typical applicant: {dict(zip(feature_names, mean_applicant.round(1)))}")
print(f"this applicant: {dict(zip(feature_names, applicant))}")
print(f"predicted default probability: {model.predict_proba([applicant])[0, 1]:.2f}")
print()
print("reason codes, ranked by how much each factor pushed risk up:")
order = np.argsort(-contributions)
for i in order:
direction = "increases" if contributions[i] > 0 else "decreases"
print(f" {feature_names[i]:>16}: {direction} risk (contribution {contributions[i]:+.2f})")typical applicant: {'income': 63.0, 'debt': 12.0, 'missed_payments': 1.4}
this applicant: {'income': 32, 'debt': 21, 'missed_payments': 3}
predicted default probability: 1.00
reason codes, ranked by how much each factor pushed risk up:
debt: increases risk (contribution +5.57)
missed_payments: increases risk (contribution +0.63)
income: increases risk (contribution +0.13)What actually happened
contributions = model.coef_[0] * (applicant - mean_applicant) decomposes the applicant's log-odds into three exact pieces, one per feature, each measured against a typical applicant rather than zero. That comparison to a typical applicant is what makes a reason meaningful — "your debt is unusually high" is checkable; "your debt coefficient is 0.615" is not.
debt dominates here, by a wide margin — this applicant's debt (21k) is far above the training data's average (12k), and the model weighs debt heavily. That becomes reason code #1: "outstanding debt relative to typical applicants."
Notice income's contribution is small, even though this applicant's income (32k) is well below average (63k). This is the same effect flagged in credit scoring: the model's coefficient on income is small because income, debt and missed payments move together in this tiny training set, so the model leaned on debt to do most of the explaining. This is exactly why real credit models are built and reviewed with correlated features in mind — a naive reason-code generator can under-credit a real factor that a correlated feature has absorbed.
- This decomposition is exact only because logistic regression is additive in log-odds. It does not carry over unchanged to a tree ensemble, where "each feature's contribution" is not naturally defined — that gap is why SHAP exists.
- Ranking by
contributions[i], notabs(contributions[i]), is deliberate: an adverse action reason should only cite factors that pushed toward decline, not factors that helped the applicant.
Common mistakes
Reporting the coefficient, not the contribution. A large coefficient on a feature that barely varies among real applicants can be a small real-world factor. Always multiply by how far this applicant sits from typical, not only the raw coefficient.
Citing more reasons than the applicant can act on. Most reason-code frameworks cap the list at three or four — the top few factors, not an exhaustive dump of every feature.
Using this exact technique on a non-linear model. Trees, forests and boosted ensembles do not decompose into additive per-feature pieces this cleanly. SHAP and LIME exist specifically to approximate this kind of explanation for models where it is not exact.
Treating the explanation as separate from the model. A reason code is only trustworthy if it is generated from the exact model that made the decision — a mismatch between the deciding model and the explaining model is itself a regulatory and correctness problem.
Try it yourself
Change applicant to [50, 14, 2] — a genuinely borderline case. Print the reason codes and notice how much closer the three contributions are to each other, compared to the clear-cut declined case above. Borderline cases are where reason-code quality matters most.
What to learn next
- SHAP and LIME — the general technique for models without a clean additive structure.
- Multicollinearity and VIF — why a correlated feature can absorb credit that belongs elsewhere.
- Model risk management — where explanation quality itself gets reviewed before deployment.
Researcher — Mathematics and papers.
Additive decompositions in general
For a linear-in-log-odds model, the decomposition used above is exact:
logit(x) = b_0 + SUM_i b_i * x_i = b_0 + SUM_i b_i * x_bar_i + SUM_i b_i * (x_i - x_bar_i)b_0— interceptb_i— coefficient for featureix_bar_i— a chosen reference value for featurei(population mean, a specific "typical approved applicant," or a regulator-approved reference profile)
The middle term is a constant baseline; the final sum is the exact per-feature attribution relative to that baseline. Reference-point choice is a modelling decision with real regulatory consequences, since it changes which factors appear to matter for a given applicant.
Shapley values as the general case
For a model f that is not additive, the standard axiomatic choice is the Shapley value from cooperative game theory, adapted to feature attribution by Lundberg & Lee (2017) as SHAP:
phi_i = SUM_{S subseteq F \ {i}} |S|! (|F| - |S| - 1)! / |F|! * [ f(S union {i}) - f(S) ]F— the full set of featuresS— a subset of features not includingif(S)— the model's expected output using only the features inS, others marginalised outphi_i— featurei's attributed contribution
This is the unique attribution satisfying four properties simultaneously: efficiency (attributions sum to the model's output minus its baseline), symmetry, dummy (a feature with no effect gets zero credit), and additivity across models. For a linear model, SHAP values reduce exactly to the decomposition shown in the developer section — which is why that section's code is a correct special case, not an approximation.
Cost
Exact Shapley values require evaluating f over all 2^|F| feature subsets, intractable beyond a handful of features. TreeSHAP (Lundberg et al., 2020) computes exact Shapley values for tree ensembles in polynomial time by exploiting tree structure directly, which is why SHAP is practically usable on gradient-boosted credit models despite the combinatorial definition.
Regulatory specifics (illustrative, jurisdiction-dependent)
In the United States, Regulation B (implementing the Equal Credit Opportunity Act) requires that adverse action notices state the "principal reason(s)" for denial in specific, non-generic terms — checklists of acceptable and unacceptable reason phrasing have existed since long before machine learning, originally written for manual scorecards. Regulatory guidance since around 2020 has explicitly grappled with whether SHAP-style explanations for complex models satisfy this standard; this remains an active compliance question, not settled law, and requirements differ across jurisdictions.
Key references
- Lundberg, S. M., & Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions. NeurIPS — the original SHAP paper.
- Lundberg, S. M., et al. (2020). From Local Explanations to Global Understanding with Explainable AI for Trees. Nature Machine Intelligence — TreeSHAP.
- Consumer Financial Protection Bureau guidance on AI/ML in credit underwriting and Regulation B adverse action requirements (consult current CFPB publications for the applicable version).
Current state
SHAP is the de facto standard for explaining tree-ensemble credit models in production, but there is no legal consensus yet on whether a SHAP value alone constitutes a compliant "specific reason," and several institutions still maintain a simpler, fully transparent scorecard purely for the customer-facing explanation. Treat this as an open regulatory-technical question, not a solved one.
What to learn next
- SHAP and LIME — the general treatment of both methods.
- Model risk management — where explanation methodology itself must be validated.
- Fairness metrics — the audit use case these explanations exist to support.