Scoping an ML Project

Getting knowledge out of a domain expert

Domain experts hold the features, labels, and failure warnings your data cannot reveal, and extracting that knowledge is a skill with concrete techniques.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A domain expert carries years of pattern knowledge your data does not show, and one hour of asking the right questions can outperform a month of model tuning.

An old mechanic hears your engine for ten seconds and says "loose belt". He is not guessing — thirty years of engines taught his ears a pattern. Ask him to write down how he knows, though, and he struggles. The knowledge is in his hands and ears, not in sentences.

Nurses, loan officers, farmers, and fraud analysts all hold knowledge in this form. Your project needs it, and they cannot hand it over unaided. Extracting it is a skill.

Why it exists

The data you have describes what happened. The expert knows why, what is missing from the tables, and which patterns are traps. Models built without this input rediscover what the team already knew, at great cost. Worse, they learn a pattern the expert could have debunked in a sentence.

The catch: asking "how do you decide?" produces textbook answers, not real ones. People describe how they believe they decide. The craft is asking questions that reach the real process.

How it works

Three techniques, in rising order of power:

1. WALK THROUGH CASES     "talk me through these 10 real examples"
   → watch what they look at FIRST — that is a feature

2. ASK FOR THE EXCEPTIONS "when does the usual rule fail?"
   → exceptions become features, guardrails, or label fixes

3. SHOW MODEL MISTAKES    "the model got these wrong — what do you see?"
   → the expert debugs your model with knowledge you lack

Concrete cases beat general questions every time. "How do you spot a risky loan?" gets a shrug. "Why did you reject this one?" gets "salary credited in cash, and the shop is rented — together, that worries me". That sentence contains two features and an interaction.

A real example you have seen

Weather apps improved when meteorologists' hand rules were turned into features. One such rule: clouds like this over the coast mean rain inland by evening. The pattern repeats in medicine, farming, and fraud: the model industrialises what an expert already half-knew.

Remember this

  • Experts hold real patterns they cannot articulate on request. Use cases, not general questions.
  • What the expert looks at first is a feature. What they call an exception is a guardrail.
  • Showing the expert your model's mistakes turns them into your best debugger.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn

Outputs verified with scikit-learn 1.7 and numpy 1.26.

One sentence from a nurse, worth 30 points of AUC

The dataset has blood pressure and age. The model sees numbers; the nurse knows both very low and very high readings are dangerous. Watch that one sentence become a feature.

expert_feature.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(1)
n = 3000

# Systolic blood pressure readings and patient age.
bp = rng.normal(125, 25, n).clip(70, 210)
age = rng.normal(55, 12, n).clip(20, 90)

# The truth the nurse knows: BOTH very low and very high readings are dangerous.
danger = (np.abs(bp - 120) + rng.normal(0, 8, n) > 30).astype(int)

X_raw = np.column_stack([bp, age])
# One line of domain knowledge, turned into a feature.
X_expert = np.column_stack([bp, age, np.abs(bp - 120)])

for name, X in [("raw readings", X_raw), ("+ expert feature", X_expert)]:
    X_tr, X_te, y_tr, y_te = train_test_split(X, danger, random_state=0)
    m = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)
    auc = roc_auc_score(y_te, m.predict_proba(X_te)[:, 1])
    print(f"{name:18s} AUC {auc:.3f}")
Output
raw readings       AUC 0.640
+ expert feature   AUC 0.948

The walkthrough

Why the raw model fails: danger here is U-shaped — high at both extremes of blood pressure. A linear model can only learn "higher is worse" or "lower is worse", never "both ends are worse". It settles for a nearly useless compromise: AUC 0.640, barely above coin-flip ranking.

The expert feature is abs(bp - 120) — distance from normal. In that transformed space the pattern is linear, and the same model jumps to 0.948. The nurse did not know the word "non-linearity". She knew the pattern, which is the part that matters.

AUC, for reading the numbers: the probability the model ranks a random dangerous case above a random safe one. 0.5 is guessing; 1.0 is perfect ranking.

A tree model would find this pattern alone — eventually, with enough data. The expert feature gets it instantly, with less data, and produces a model whose reasoning the expert recognises. That last property pays off again at explanation time.

Common mistakes

Interviewing with abstractions. "What factors matter?" retrieves the textbook. Ten printed real cases on the table retrieves the experience. Always cases.

Dismissing rules that sound unscientific. "Applications submitted at 3 am worry me" sounds like superstition and is a testable hypothesis, one groupby away. Test the expert's hunches before filing them under folklore — they are hypotheses from a person with years of labelled data in their head.

Encoding beliefs without validation. The reverse error: experts are sometimes wrong, and confidently so. Every expert feature earns its place the same way — measured lift on held-out data, like the AUC comparison above.

One interview at the start, then silence. The highest-value session is after the first model, reviewing its errors together. Error analysis with an expert in the room routinely finds label problems and missing features in an hour.

Try it yourself

Delete age from both feature sets and re-run — notice how little it was contributing. Then try giving the raw model a fighting chance without the expert: add bp**2 as a feature. Quadratic terms can express U-shapes; check how close it gets to 0.948, and think about which version you could defend to the nurse.

What to learn next

Researcher — Mathematics and papers.

The knowledge acquisition bottleneck

Expert-systems research named this problem in the 1980s: extracting rules from experts was the binding constraint on building MYCIN-style systems (Feigenbaum, 1977; Buchanan & Shortliffe, 1984, Rule-Based Expert Systems). Two findings from that literature survive directly into modern ML practice:

  1. Experts' verbal reports are unreliable models of their own process (Nisbett & Wilson, 1977, Telling More Than We Can Know, Psych. Review). Hence case-based elicitation over introspection.
  2. Expertise is pattern recognition over a large stored case base (Klein's recognition-primed decision model, 1993) — which is why showing concrete instances retrieves knowledge that questions cannot.

Formal channels for injecting domain knowledge

Feature construction is one channel among several:

  • Feature transforms — the developer example; knowledge as coordinates in which the hypothesis class is well-specified.
  • Constraints — monotonicity (loan risk non-increasing in income) enforced in gradient boosting and lattice models (Gupta et al., 2016, Monotonic Calibrated Interpolated Look-Up Tables, JMLR). A monotonicity constraint is an expert statement with regulariser semantics.
  • Priors — Bayesian formulations encode expert belief as $p(\theta)$; elicitation methodology is its own field (O'Hagan et al., 2006, Uncertain Judgements).
  • Weak supervision — experts write labelling functions rather than labels; Snorkel (Ratner et al., 2017, VLDB) denoises and combines them, scaling one expert's rules to millions of labels.
  • Shape functions review — with GAMs (generalized additive models), experts can read each learned univariate effect and veto artefacts; Caruana et al. (2015), Intelligible Models for HealthCare (KDD), report the canonical case: a pneumonia model learned "asthma lowers risk" — an artefact of aggressive treatment, caught precisely because the model was inspectable by clinicians.

Elicitation bias

Expert input inherits human judgement biases — availability, recency, anchoring (Tversky & Kahneman, 1974). Structured protocols mitigate: multiple experts elicited independently before discussion (Delphi method), calibration questions with known answers to weight experts (Cooke, 1991, Experts in Uncertainty), and always the final arbiter: held-out validation of every encoded belief. The asymmetry worth remembering — a wrong expert feature costs a validation experiment; a missing expert feature can cost the project.

What to learn next