Scoping an ML Project

Explaining a model to the person who decides

Decision-makers do not need to understand the model — they need its behaviour translated into consequences, costs, and reasons they can act on.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Explaining a model means translating its behaviour into decisions and consequences — not describing its machinery.

A good doctor does not explain your blood report by describing the testing machine. She says: "this number is high, it means risk, here is what we change." The machine is her concern. The meaning and the action are yours.

You are the doctor in the model conversation. The stakeholder — the manager, the regulator, the loan officer — needs meaning and action, in their language.

Why it exists

Projects die in the approval meeting, not the notebook. A technically excellent model that the decision-maker cannot trust does not get deployed, and mistrust usually starts with an explanation full of machinery: architectures, hyperparameters, AUC.

The stakeholder has three real questions, and every explanation should be built from them:

  • What will change in my process?
  • How often is it wrong, and what does a mistake cost me?
  • When it decides something, can I see why?

How it works

Translate each model fact into a decision fact:

model fact                    →  decision fact
"precision 0.83 at our cut-off"  →  "of 100 flagged, about 83 are real problems"
"recall 0.71"                    →  "we catch about 7 of every 10 cases"
"missed payments matters most"   →  "missed payments lower approval odds"
"probability 0.08"               →  "we'd approve 8 of 100 applicants like this"

Two habits carry most of the weight. First, use frequencies, not probabilities — "8 of 100 people like this" lands where "0.08" does not. Second, always name the mistakes: a stakeholder who learns the error rate from you will trust you; one who discovers it in production will not.

A real example you have seen

Weather apps went from "60% chance of rain" to "rain likely around 4 pm — carry an umbrella". Same model underneath. The second version translates a probability into a decision, and that translation is the entire difference in usefulness.

Remember this

  • Explain consequences, never machinery. Nobody deciding a budget needs the word "hyperparameter".
  • Frequencies beat probabilities: "8 of 100", not "0.08".
  • Volunteer the error rate and its cost. Trust comes from named weaknesses.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26.

Turning coefficients into sentences

A loan model, explained twice: once as what it learned overall, once for a single applicant. Both outputs are meant to be pasted into a slide, unedited.

decision_language.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(5)
n = 1200
income = rng.normal(50000, 15000, n).clip(12000, 120000)
existing_loans = rng.integers(0, 5, n)
missed_payments = rng.poisson(0.8, n).clip(0, 6)

logit = 0.00004 * income - 0.5 * existing_loans - 0.9 * missed_payments - 0.3
repaid = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int)

names = ["income", "existing loans", "missed payments"]
X = np.column_stack([income, existing_loans, missed_payments])
scaler = StandardScaler().fit(X)
model = LogisticRegression().fit(scaler.transform(X), repaid)

print("what the model learned, in decision language:\n")
for name, coef in zip(names, model.coef_[0]):
    direction = "raises" if coef > 0 else "lowers"
    print(f"  a typical step up in {name:16s} {direction} the odds of "
          f"repayment by {abs(np.exp(coef) - 1):.0%}")

# One applicant, explained.
applicant = np.array([[28000, 3, 2]])
z = scaler.transform(applicant)[0]
contributions = model.coef_[0] * z
prob = model.predict_proba(scaler.transform(applicant))[0, 1]
print(f"\napplicant: income 28000, 3 loans, 2 missed payments")
print(f"predicted chance of repayment: {prob:.0%}")
worst = np.argsort(contributions)[:2]
print("biggest factors pulling the score down:")
for i in worst:
    print(f"  {names[i]} (this applicant: {applicant[0][i]:.0f})")
Output
what the model learned, in decision language:

  a typical step up in income           raises the odds of repayment by 72%
  a typical step up in existing loans   lowers the odds of repayment by 49%
  a typical step up in missed payments  lowers the odds of repayment by 54%

applicant: income 28000, 3 loans, 2 missed payments
predicted chance of repayment: 8%
biggest factors pulling the score down:
  missed payments (this applicant: 2)
  income (this applicant: 28000)

The walkthrough

"A typical step up" is doing careful work. Features were standardised, so each coefficient describes one standard deviation of change — a typical step, in plain words. Without scaling, "one unit of income" means one rupee and every sentence becomes nonsense.

np.exp(coef) converts to odds ratios. Logistic coefficients live in log-odds, which no human thinks in. Exponentiating gives "multiplies the odds by 1.72", rendered above as "raises the odds by 72%" — defensible and readable.

The per-applicant block answers the regulator's question. "Why was this person scored low?" gets specific factors with the applicant's own values. This is the model's side of the human-in-the-loop screen.

This transparency is why linear models still win regulated deals. For black-box models, tools like SHAP reconstruct similar statements approximately; a logistic model gives them exactly.

Common mistakes

Leading with the ROC curve. The stakeholder nods, remembers nothing, trusts less. Lead with the decision table: flagged per day, caught, missed, cost of each.

Odds versus probability, blurred. "Raises the odds by 72%" is not "raises the probability by 72 points". Odds ratios are the honest quote for logistic models; when the audience cannot carry odds, switch to frequencies ("8 of 100 applicants like this").

Implying cause. "Missed payments lower the score" describes the model. "Paying on time will get you approved" is a causal promise the model never made. Keep the verbs descriptive, especially in writing.

Hiding the failure modes. Volunteer where the model is weakest — thin-file applicants, a bad slice — with numbers. Slice-based evaluation generates exactly this content, and presenting it is what separates trusted teams from oversold ones.

Try it yourself

Write the rejection sentence a bank clerk could read to this applicant, using only the output above and no ML words. Under 40 words. Harder than any part of the modelling — and this sentence is what the customer experiences as "the AI".

What to learn next

Researcher — Mathematics and papers.

Odds ratios and their communication limits

For logistic regression, $\beta_j$ is the change in log-odds per unit of $x_j$; $e^{\beta_j}$ is the odds ratio:

$$ \log \frac{p}{1-p} = \beta_0 + \sum_j \beta_j x_j \qquad \frac{\text{odds}(x_j + 1)}{\text{odds}(x_j)} = e^{\beta_j} $$

Where $p$ is the predicted probability and odds $= p/(1-p)$. Odds ratios are constant across the feature range; probability changes are not — the same $e^{\beta}$ moves 50%→64% but 5%→8%. Quoting probability deltas therefore requires a reference point; quoting odds ratios requires an audience that understands odds. There is no free option, which is the technical root of the beginner block's advice.

Frequency formats: the empirical case

Gigerenzer and Hoffrage (1995), How to Improve Bayesian Reasoning Without Instruction (Psych. Review), showed diagnostic-inference accuracy in physicians rises dramatically when information arrives as natural frequencies ("8 of 1000") instead of conditional probabilities. Follow-up work (Gigerenzer et al., 2007, Helping Doctors and Patients Make Sense of Health Statistics) extends this to risk communication generally. For ML: confusion matrices at the operating threshold, stated per 100 or 1000 cases, are the evidence-backed format.

Post-hoc explanation methods and their caveats

For non-linear models, per-decision attributions come from:

  • LIME (Ribeiro et al., 2016, KDD) — local surrogate models; unstable under resampling (Alvarez-Melis & Jaakkola, 2018).
  • SHAP (Lundberg & Lee, 2017, NeurIPS) — Shapley-value attributions with consistency axioms; exact TreeSHAP for tree ensembles (Lundberg et al., 2020, Nature MI).
  • Caveats that matter in stakeholder settings: attributions explain the model, not the world; correlated features split credit arbitrarily; and adversarial constructions can make biased models produce innocuous explanations (Slack et al., 2020, Fooling LIME and SHAP, AIES). An explanation pipeline is part of the audited surface, not a decoration.

Regulatory context

Adverse-action requirements (US ECOA/Reg B) require specific principal reasons for credit denial — the per-applicant block above is shaped by that requirement. Counterfactual explanations (Wachter et al., 2017, Counterfactual Explanations Without Opening the Black Box, Harvard JOLT) propose "smallest change flipping the decision" as the legally useful unit; recourse formulations (Ustun et al., 2019, Actionable Recourse in Linear Classification, FAT*) add actionability constraints — income can rise, age cannot reverse. The open problem: recourse statements drift toward causal promises, which predictive models cannot honour; see the "implying cause" mistake above.

What to learn next