SHAP and LIME
SHAP splits a single prediction into a fair share per feature; LIME fits a small readable model near one point — here both are built from scratch.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
SHAP and LIME both answer one question: why did the model say that about this one case?
The analogy you have already lived
Four friends share an auto to the station and the fare is 300 rupees. Splitting it four ways feels wrong, because two of them got off halfway.
So you work out what the fare would have been for every possible combination of passengers. Then you give each person the average extra cost they added by being in the auto. That is a fair share.
SHAP does exactly this with features instead of passengers. The prediction is the fare. Each feature gets the average amount it added, across every combination of the other features.
And LIME is a different trick
LIME takes a different route to the same goal.
Imagine a winding mountain road. Over the whole journey it curves in every direction, and no straight line describes it. But stand at one bend and look at the two metres in front of you, and it is basically straight.
LIME zooms in until the complicated model looks like a straight line. Then it reports the straight line, because a straight line is readable.
SHAP asks "what is each feature's fair share?" LIME asks "what does the model look like right here?"
Why these were invented
A model can tell a bank that your loan is declined. The bank has to tell you why. "The gradient boosting ensemble decided" helps nobody.
What is needed is a sentence: your application was declined mainly because of four late payments last year, and helped a little by your income. That is a local explanation — an explanation of one decision, not of the model in general.
How it looks
average prediction for everyone → 7
|
income is high → pushes up +18
debt is high → pushes down -19
age → barely moves -1
4 late payments → pushes down -22
|
this person's prediction → -17That column of numbers adds up. The starting point plus every feature's push lands exactly on the final answer. That adding-up property is what makes SHAP trustworthy, and it is why banks and hospitals prefer it.
What is honestly hard here
Both methods are approximations that need a reference point — a notion of what "normal" looks like. Change the reference and every number changes.
And both can be fooled. Researchers have built models that discriminate on real data while showing innocent-looking explanations. An explanation is evidence, not proof.
This part is confusing for almost everyone the first time. Read it twice — that is normal.
Remember this
- Both explain one prediction, not the whole model.
- SHAP gives each feature a fair share that adds up to the prediction.
- LIME fits a straight line near one point and reports its slopes.
What to learn next
- Privacy in machine learning — what a model leaks about the people in its training set.
- Model cards and documentation — where explanations belong in a release.
- XGBoost — the tree models TreeSHAP was built for.
Developer — Code and libraries.
Build both from scratch first
The libraries are excellent and you should use them in production. But shap and lime are much easier to use correctly once you have written the twenty lines underneath them. Everything below runs on NumPy and scikit-learn alone.
Setup
pip install numpy scikit-learnFor production use later: pip install shap lime. Nothing on this page needs them.
Part 1 — exact Shapley values, all sixteen coalitions
With four features there are only sixteen subsets, so we can compute the exact answer rather than sampling. The exact form is the definition that KernelSHAP approximates.
import itertools, math
import numpy as np
from sklearn.ensemble import GradientBoostingRegressor
rng = np.random.default_rng(5)
n, names = 800, ["income", "debt", "age", "late_payments"]
X = np.c_[rng.normal(50, 15, n), rng.normal(20, 8, n),
rng.normal(35, 10, n), rng.poisson(1.2, n).astype(float)]
# a non-linear truth, so a linear model would not capture it
y = 0.8 * X[:, 0] - 1.5 * X[:, 1] - 6.0 * X[:, 3] ** 1.3 + 0.2 * X[:, 2] + rng.normal(0, 3, n)
model = GradientBoostingRegressor(random_state=0).fit(X, y)
background = X[:200] # the "what if we knew nothing" reference set
base_value = model.predict(background).mean()
x = np.array([70.0, 35.0, 29.0, 4.0]) # one applicant we must explain
fx = float(model.predict(x[None, :])[0])
def v(subset):
"""Average prediction when the features in `subset` are held at x's values
and every other feature is drawn from the background data."""
Z = background.copy()
for i in subset:
Z[:, i] = x[i]
return float(model.predict(Z).mean())
d = len(names)
shap = np.zeros(d)
for i in range(d):
others = [j for j in range(d) if j != i]
for r in range(d): # all subsets S not containing i
for S in itertools.combinations(others, r):
w = math.factorial(r) * math.factorial(d - r - 1) / math.factorial(d)
shap[i] += w * (v(tuple(S) + (i,)) - v(S))
print("EXACT SHAPLEY VALUES (all 16 coalitions, no sampling)")
for nm, s in zip(names, shap):
print(f" {nm:14s} {s:+9.2f}")
print(f"\n base value (average prediction) : {base_value:9.2f}")
print(f" + sum of shapley values : {shap.sum():+9.2f}")
print(f" = reconstructed prediction : {base_value + shap.sum():9.2f}")
print(f" actual model prediction : {fx:9.2f}")
# ---- LIME: fit a small weighted linear model in the neighbourhood of x ----
from sklearn.linear_model import Ridge
m = 2000
mask = rng.random((m, d)) < 0.5 # 1 = keep x's value, 0 = replace it
donors = background[rng.integers(0, len(background), m)]
Z = np.where(mask, x, donors)
preds = model.predict(Z)
# points that keep more of x's real values count more
dist = np.sqrt(((~mask).sum(axis=1)) / d)
weights = np.exp(-(dist ** 2) / 0.25 ** 2)
lime = Ridge(alpha=1.0).fit(mask.astype(float), preds, sample_weight=weights)
print("\nLIME COEFFICIENTS (local linear fit, 2000 perturbations)")
for nm, c in zip(names, lime.coef_):
print(f" {nm:14s} {c:+9.2f}")
print(f" local fit R2 on the weighted sample: "
f"{lime.score(mask.astype(float), preds, sample_weight=weights):.3f}")EXACT SHAPLEY VALUES (all 16 coalitions, no sampling)
income +18.24
debt -19.22
age -1.02
late_payments -22.15
base value (average prediction) : 7.01
+ sum of shapley values : -24.15
= reconstructed prediction : -17.14
actual model prediction : -17.14
LIME COEFFICIENTS (local linear fit, 2000 perturbations)
income +13.93
debt -13.96
age -0.21
late_payments -14.87
local fit R2 on the weighted sample: 0.714The three things in that output that matter
Additivity is exact, not approximate. 7.01 + (-24.15) = -17.14, and the model predicted -17.14. This is the efficiency property, and it holds to floating-point precision because we enumerated every coalition instead of sampling them. When a SHAP library reports values that do not sum to the prediction, sampling error is the reason.
The rankings agree; the magnitudes do not. Both methods put late_payments first, debt second, income third, age last. But SHAP says -22.15 for late payments and LIME says -14.87. They are not measuring the same quantity, and quoting them interchangeably is wrong.
LIME's R2 of 0.714 is the number nobody prints. It says the local straight line explains 71% of the model's behaviour in that neighbourhood. The other 29% is explanation error. If you report LIME coefficients without this number, you are hiding the error bar on your own explanation.
Line by line
v(subset) is the value function. Holding a feature at x's value and drawing the rest from background computes an interventional (marginal) expectation. The alternative, conditional expectation, respects feature correlations but needs a model of the data distribution. The two disagree whenever features are correlated, and libraries differ in which they implement.
w = r! (d - r - 1)! / d! is the Shapley weight — the probability that subset S is exactly the set of features arriving before feature i in a random ordering. Summing over all subsets is the same as averaging over all d! orderings.
mask = rng.random((m, d)) < 0.5 builds LIME's interpretable representation: a binary vector saying which features kept their real value. The linear model is fitted on those binary flags, not on the raw feature values, which is why the coefficients read as "effect of having this feature's real value".
weights = np.exp(-(dist ** 2) / 0.25 ** 2) is the locality kernel. That 0.25 is the kernel width, and it is the most consequential free parameter in LIME. Widen it and the explanation becomes global and inaccurate; narrow it and the fit becomes unstable.
Cost, before you run this on real data
Exact Shapley is $O(2^d)$ model evaluations per explanation. Four features means 16 coalitions. Twenty features means over a million, per row explained.
That is why real tools do not do this. shap.TreeExplainer computes exact tree SHAP in polynomial time (Lundberg et al., 2020) and is the right choice for any tree ensemble. KernelExplainer samples coalitions and is model-agnostic and slow. Deep models use gradient-based approximations.
Common mistakes
Using the whole training set as background. Cost scales linearly in background size. Use a k-means summary of 100 rows, and state which background you used, because the numbers are meaningless without it.
Reading SHAP values as causal effects. A value of +18.24 for income means the model's output moves that much. It does not mean raising this person's income by any amount changes anything in the world.
Averaging absolute SHAP values and calling it global importance. It is a reasonable summary, and it is not the same as permutation importance. Positive and negative effects cancel differently under each.
Explaining correlated features independently. The interventional value function builds records with an income of 70 and a debt drawn from someone else's row. Some of those people cannot exist. Explanations of impossible people are not reliable.
Running LIME once and trusting it. It samples. Run it five times with different seeds; if the ranking moves, raise m or report the instability.
Try it yourself
Change background = X[:200] to background = X[:200][X[:200, 3] >= 2] so the reference group is people who already have late payments. The late_payments contribution should shrink sharply — not because the applicant changed, but because "normal" did. That sensitivity is the honest headline of both methods.
What to learn next
- Privacy in machine learning — what a model leaks about the people in its training set.
- Model cards and documentation — where explanations belong in a release.
- XGBoost — the tree models TreeSHAP was built for.
Researcher — Mathematics and papers.
Shapley values
The Shapley value (Shapley, 1953) is the unique allocation satisfying efficiency, symmetry, dummy and additivity. For a set function $v: 2^{[d]} \to \mathbb{R}$,
$$ \phi_i(v) = \sum_{S \subseteq [d] \setminus {i}} \frac{|S|!\,(d - |S| - 1)!}{d!}\Big(v(S \cup {i}) - v(S)\Big) $$
- $[d]$ — the index set of features; $i$ the feature being credited.
- $S$ — a coalition of features excluding $i$.
- $v(S)$ — the value of coalition $S$ (below, an expected model output).
- The combinatorial factor is the probability that $S$ precedes $i$ in a uniformly random permutation.
Efficiency gives $\sum_i \phi_i = v([d]) - v(\emptyset)$, which is the additivity checked numerically above.
The value function is a modelling choice
SHAP (Lundberg and Lee, 2017) instantiates $v$ with an expected model output. Two variants exist and are routinely confused.
$$ v_{\text{int}}(S) = \mathbb{E}{X{\bar S}}\big[f(x_S, X_{\bar S})\big], \qquad v_{\text{cond}}(S) = \mathbb{E}\big[f(X) \mid X_S = x_S\big] $$
$\bar S$ is the complement of $S$, $x_S$ the explained instance restricted to $S$, and $X_{\bar S}$ a random draw of the remaining features.
The interventional form breaks feature dependence and evaluates $f$ off-manifold. The conditional form stays on-manifold and spreads credit to correlated features that the model never reads. Chen et al. (2020), True to the Model or True to the Data?, argue the choice follows from the question: interventional for auditing the model's mechanism, conditional for reasoning about the data-generating process. Janzing et al. (2020) argue from a causal standpoint that the interventional form is the correct default for explaining $f$.
KernelSHAP
Lundberg and Lee show that Shapley values are recovered as the solution of a weighted least-squares problem over the $2^d$ coalitions, with the Shapley kernel
$$ \pi(S) = \frac{d - 1}{\binom{d}{|S|}\,|S|\,(d - |S|)} $$
The kernel diverges at $|S| \in {0, d}$, which are handled as hard constraints. This is exactly LIME's optimisation problem with a specific kernel, a specific loss and a specific regulariser — the unification result of the paper. Sampling coalitions from $\pi$ gives an unbiased but high-variance estimator; convergence is slow in $d$.
TreeSHAP and its subtleties
Lundberg et al. (2020), From Local Explanations to Global Understanding with Explainable AI for Trees (Nature MI), give an $O(TLD^2)$ algorithm for tree ensembles — $T$ trees, $L$ leaves, $D$ depth — versus exponential enumeration. The default tree_path_dependent mode uses the training-set coverage counts stored in the tree as an implicit conditional distribution. That is neither $v_{\text{int}}$ nor a clean $v_{\text{cond}}$, and it can assign non-zero attribution to features the model does not use, as analysed by Sundararajan and Najmi (2020). Pass an explicit background dataset for interventional values when the distinction matters.
LIME
Ribeiro et al. (2016) solve
$$ \xi(x) = \arg\min_{g \in G} \; \mathcal{L}\big(f, g, \pi_x\big) + \Omega(g) $$
with $G$ a class of interpretable models (sparse linear), $\pi_x$ a proximity kernel around $x$, $\mathcal{L}$ a locality-weighted squared loss and $\Omega$ a complexity penalty (typically K-LASSO to a fixed feature count).
Known weaknesses, all reproducible:
- Instability. Alvarez-Melis and Jaakkola (2018) show large explanation changes for small input changes. Zafar and Khan (2019) propose deterministic sampling to address it.
- Kernel width sensitivity. No principled selection procedure exists; the default heuristic is $\sqrt{d} \times 0.75$ on scaled data, chosen empirically.
- Off-manifold sampling. Perturbations are drawn independently per feature, so correlated data produces implausible neighbours.
The adversarial result
Slack et al. (2020) construct a scaffolding attack: a classifier that detects whether an input came from the real data distribution or from LIME/SHAP's perturbation distribution, behaves discriminatorily on the former and innocuously on the latter. Both explainers then report benign attributions for a model that is biased in deployment.
The defence is not a better attribution method. It is on-manifold perturbation, plus treating explanations as one input to an audit rather than its conclusion.
Axiomatic alternatives
Integrated Gradients (Sundararajan et al., 2017) satisfies completeness and implementation invariance for differentiable models at the cost of a baseline choice and $O(m)$ gradient evaluations along a path. Aumann-Shapley is its cooperative-game analogue. For deep networks it is usually the cheaper defensible option; see SHAP's DeepExplainer for the hybrid.
Papers
- Lundberg and Lee, A Unified Approach to Interpreting Model Predictions, NeurIPS 2017 — arxiv.org/abs/1705.07874
- Ribeiro, Singh and Guestrin, "Why Should I Trust You?", KDD 2016 — arxiv.org/abs/1602.04938
- Lundberg et al., Explainable AI for Trees, Nature Machine Intelligence 2020
- Chen et al., True to the Model or True to the Data?, 2020 — arxiv.org/abs/2006.16234
- Sundararajan and Najmi, The Many Shapley Values for Model Explanation, ICML 2020 — arxiv.org/abs/1908.08474
- Sundararajan, Taly and Yan, Axiomatic Attribution for Deep Networks, ICML 2017 — arxiv.org/abs/1703.01365
- Slack et al., Fooling LIME and SHAP, AIES 2020 — arxiv.org/abs/1911.02508
What to learn next
- Privacy in machine learning — what a model leaks about the people in its training set.
- Model cards and documentation — where explanations belong in a release.
- XGBoost — the tree models TreeSHAP was built for.