AI Safety and Ethics

Explainability

Explainability is the engineering of checkable answers to "why did the model say that" — and different methods answer different questions.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why this matters more than it sounds
  4. The classic story worth remembering
  5. The two kinds of "why"
  6. What is honestly hard here
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Explainability is getting a model to show its reasons in a form a person can check.

The analogy you have already lived

Think of a school exam where you had to show your working. The final answer alone was worth little. The steps were worth most of the marks.

There was a good reason for that. A right answer with wrong working means you got lucky, and you will get the next one wrong. Working lets the teacher find the exact place you went astray.

A model that gives an answer with no working is a student who refuses to show steps. It might be right. You cannot tell, and you cannot fix it.

Why this matters more than it sounds

You need reasons for four practical jobs, and none of them are philosophical.

Debugging. Your model rejects an application. Was it the income, or was it a stray column that leaked the answer? Only reasons tell you.

Trust from users. A doctor will not act on "the machine said so". A reason turns an output into evidence a professional can weigh.

Rights. In several countries, a person affected by an automated decision can ask why. "The network decided" is not an answer that survives a complaint.

Catching cheating models. This is the big one. A model can be right for a reason that will stop working tomorrow.

The classic story worth remembering

Researchers built a model that told wolves from huskies with high accuracy. It looked excellent.

Then they looked at what it was using. The wolf photos had snow in the background. The husky photos did not. The model had learned to detect snow.

It scored well on every test they had. It would fail the first time somebody photographed a husky in Manali.

A model can be accurate and wrong at the same time. Only explanations catch that, and no accuracy number ever will.

The two kinds of "why"

   GLOBAL  →  "What does this model use, in general?"
              income matters a lot, pin code matters a little

   LOCAL   →  "Why THIS person, THIS time?"
              your application was declined mainly because of
              the 4 missed payments last year

Global answers are for the team building the model. Local answers are for the person affected by it. You need both, and they are computed differently.

What is honestly hard here

An explanation is a story told about a model. It is not the model.

Two explanation methods can look at the same model and rank the same features differently. Both can be technically correct. They are answering different questions. The developer block below shows exactly that happening, with numbers.

That is unsettling the first time you see it. It is normal, it is well known, and the fix is to know which question you asked.

Remember this

  • Explainability is about checkable reasons, not comfort.
  • Global explanations describe the model; local explanations describe one decision.
  • An explanation is a story about the model, and different methods tell different stories.

What to learn next

Developer — Code and libraries.

Three methods, one model, three different answers

The fastest way to understand explainability is to watch three standard importance methods disagree about the same model. They disagree because they ask different questions.

  • Impurity importance — how much did this column reduce error at the split points during training?
  • Permutation importance — how much does held-out performance drop when this column is shuffled?
  • Drop-column importance — how much worse is a model retrained without this column?

Setup

bash
pip install numpy scikit-learn

CPU only, a few seconds to run, no downloads.

The experiment

The data has a known ground truth: income drives the outcome, age contributes a little, and noise is junk. There is also income_copy, a near-duplicate of income — which is what happens in real pipelines when the same fact arrives from two systems.

importance_disagree.py
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.inspection import permutation_importance
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(11)
n = 1500

income   = rng.normal(0, 1, n)
income_2 = income + rng.normal(0, 0.05, n)   # a near-duplicate column: same fact, two systems
age      = rng.normal(0, 1, n)
noise    = rng.normal(0, 1, n)               # pure junk, included on purpose

# The truth: income drives the outcome, age helps a little, noise never matters.
y = 3.0 * income + 0.5 * age + rng.normal(0, 0.3, n)

X = np.c_[income, income_2, age, noise]
names = ["income", "income_copy", "age", "noise"]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=0)


def fit(Xa, ya):
    return RandomForestRegressor(n_estimators=200, random_state=0).fit(Xa, ya)


model = fit(Xtr, ytr)
print(f"held-out R2 of the full model: {model.score(Xte, yte):.4f}\n")

print("BUILT-IN IMPURITY IMPORTANCE (the number most tutorials print)")
for nm, v in zip(names, model.feature_importances_):
    print(f"  {nm:12s} {v:.3f}")

r = permutation_importance(model, Xte, yte, n_repeats=20, random_state=0, scoring="r2")
print("\nPERMUTATION IMPORTANCE (drop in held-out R2 when one column is shuffled)")
for nm, m, s in zip(names, r.importances_mean, r.importances_std):
    print(f"  {nm:12s} {m:.3f}  +/- {s:.3f}")

print("\nDROP-COLUMN IMPORTANCE (retrain from scratch without the column)")
for i, nm in enumerate(names):
    s = fit(np.delete(Xtr, i, 1), ytr).score(np.delete(Xte, i, 1), yte)
    print(f"  without {nm:12s} held-out R2 {s:.4f}")
Output
held-out R2 of the full model: 0.9855

BUILT-IN IMPURITY IMPORTANCE (the number most tutorials print)
  income       0.765
  income_copy  0.208
  age          0.024
  noise        0.002

PERMUTATION IMPORTANCE (drop in held-out R2 when one column is shuffled)
  income       1.215  +/- 0.064
  income_copy  0.116  +/- 0.006
  age          0.039  +/- 0.002
  noise        -0.000  +/- 0.000

DROP-COLUMN IMPORTANCE (retrain from scratch without the column)
  without income       held-out R2 0.9829
  without income_copy  held-out R2 0.9852
  without age          held-out R2 0.9589
  without noise        held-out R2 0.9859

Now compare the three rankings

Impurity and permutation both split income in two. income and income_copy carry identical information. The forest picked one at random at each split, so the credit for one real cause got divided across two columns. If you had reported "income importance is 0.765", you would have understated it.

Drop-column tells a completely different story. Remove income and held-out R2 falls only from 0.9855 to 0.9829 — because income_copy steps in and does the same work. Remove age and R2 falls to 0.9589, ten times further.

By drop-column, age is the most important feature. By the other two methods, income is, by a wide margin. Same model, same data, opposite conclusion.

Neither is a bug. Drop-column answers "could I do without this column", and the answer for a duplicated column is yes. Permutation answers "does this model rely on this column", and the answer is also yes. Those are different questions.

Removing noise slightly improved the score — 0.9859 against 0.9855. A junk column costs a little accuracy by giving the trees somewhere useless to split. A near-zero importance with a tiny standard deviation is how you spot it.

Line by line

permutation_importance(model, Xte, yte, ...) runs on the held-out set. Running it on training data measures memorisation. This is the single most common misuse of the function.

n_repeats=20 matters because each repeat shuffles differently. The +/- column is the standard deviation across repeats. Any importance smaller than a couple of its own standard deviations is not distinguishable from zero.

model.feature_importances_ is computed during training on the training data, and is biased toward high-cardinality and continuous features. Strobl et al. documented this in 2007. Treat it as a rough hint, never as a finding.

fit() is re-called inside the drop-column loop because drop-column importance requires actual retraining. That is why nobody runs it on a model that takes a day to train.

Common mistakes

Reporting impurity importance in a document a regulator will read. It is the cheapest and least defensible of the three. Permutation importance on held-out data is the minimum bar.

Explaining correlated features independently. Any method that perturbs one column at a time creates impossible records — an income of 200,000 next to an income_copy of 30,000. The model is being asked about a person who cannot exist. Group correlated columns and permute the group.

Confusing importance with causation. These numbers describe what the model uses. They say nothing about what would happen in the world if you changed the feature.

Using a global explanation to answer an individual's complaint. "Income is the most important feature overall" does not explain one rejection. That needs a local method — the next lesson, SHAP and LIME, covers those.

Trusting an explanation of a model you have not validated. Explaining a model that scores badly tells you why it is wrong, which is useful, but it is not evidence that it is right.

Try it yourself

Change income_2 to rng.normal(0, 1, n) so it is independent junk rather than a copy. Rerun and watch the two rankings snap back into agreement. The disagreement was caused by correlation, and correlation is the normal state of real data.

What to learn next

Researcher — Mathematics and papers.

Two families, and why the distinction is load-bearing

Interpretable models are constrained so the model is the explanation: sparse linear models, short decision lists, generalised additive models, decision sets. The explanation is exact by construction.

Post-hoc explanations are separate models fitted to describe an unconstrained model. They are approximations with an error term that is rarely reported.

Rudin (2019), Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead, makes the strong argument: on structured tabular data with meaningful features, the accuracy cost of an interpretable model is frequently negligible, and a post-hoc explanation of a black box provides the appearance of accountability without the substance. The empirical claim — that the accuracy gap is small on tabular data — has held up reasonably well; the policy conclusion remains contested.

What "explanation" formally means

There is no single definition, and the ambiguity causes most of the confusion in the literature. Three distinct targets:

Attribution. Assign each input feature a share of the output. Requires choosing a baseline $x'$ against which "contribution" is measured, and the choice materially changes the result. Additive attribution methods satisfy $f(x) - f(x') = \sum_i \phi_i$.

Counterfactual. Report the minimal change to $x$ that flips the decision. Wachter et al. (2017) solve

$$ \arg\min_{x'} \; \big(f(x') - y'\big)^2 + \lambda\, d(x, x') $$

where $x'$ is the counterfactual point, $y'$ the desired output, $d$ a distance capturing plausibility and actionability, and $\lambda$ the trade-off weight. This is the form most useful to an affected individual, because it is actionable, and the form most easily gamed, because it is actionable.

Mechanistic. Identify the internal computation. Circuit-level analysis of transformers — Elhage et al. (2021), A Mathematical Framework for Transformer Circuits; Olsson et al. (2022) on induction heads — reverse-engineers specific behaviours rather than attributing to input features. Far more expensive, and the only family that makes falsifiable claims about the mechanism.

Attribution methods are not reliable by default

Several negative results deserve to be better known.

Adebayo et al. (2018), Sanity Checks for Saliency Maps, show that several popular saliency methods produce visually similar maps when model weights are randomised. A method that is insensitive to the model it purports to explain is not explaining that model. The randomisation test they propose is cheap and should be run before any saliency method is trusted.

Kindermans et al. (2017) show that attributions from several methods change under a constant shift of the input, despite the model's behaviour being unchanged — an unreasonable sensitivity to an arbitrary baseline.

Hooker et al. (2019), ROAR, benchmark attribution methods by removing the top-ranked features and retraining. Several widely used methods performed no better than a random ranking under this test.

Slack et al. (2020), Fooling LIME and SHAP, construct a classifier that behaves in a discriminatory way on the real data distribution while producing innocuous explanations, by detecting the off-manifold perturbations both methods use. Adversarial explanation is feasible, which matters if explanations become a compliance artefact.

Faithfulness versus plausibility

The distinction (Jacovi and Goldberg, 2020) is the one to hold onto.

  • Faithfulness — does the explanation reflect the model's actual computation?
  • Plausibility — does it look convincing to a human?

Human-subject evaluation optimises plausibility. Attribution methods tuned on human ratings can become less faithful while scoring better. Faithfulness has to be measured mechanically: deletion and insertion curves, ROAR-style retraining, sufficiency and comprehensiveness on token subsets (DeYoung et al., 2020, ERASER).

Correlated features

The permutation family evaluates $f$ at points off the data manifold when features are dependent. Hooker, Mentch and Zhou (2021) analyse the resulting extrapolation bias and recommend conditional variants. Conditional permutation preserves the joint distribution but changes the question being asked — from "does the model use this feature" to "does this feature carry information beyond the others". Both are legitimate; conflating them is not.

Papers

What to learn next