When rules beat machine learning
Machine learning is one tool among several, and the first scoping decision is whether a plain written rule would solve the problem better.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Before building a model, check whether a plain written rule solves the problem — often it does, better and cheaper.
Think about separating stones from rice. You could train a child for weeks to spot stones by eye. Or you could pour everything through a sieve, because stones are bigger than grains. When a sieve works, training the child is a waste of everyone's time.
Machine learning is the trained child. A rule — a fixed instruction like "flag any transfer above ₹9,000 after 11 pm" — is the sieve.
Why this question comes first
Machine learning is expensive in ways that are easy to forget. It needs labelled examples, someone to retrain it, and someone to notice when it quietly goes wrong. A rule needs one line of code and a comment.
So the first scoping question is not "which model?". It is "does the pattern even need to be learned?".
How to decide
Ask these four questions, in order:
Can a person write the rule down? → yes → use the rule
Does the rule fit on one page? → yes → use the rule
Does the pattern shift every month? → yes → ML starts to earn its keep
Are there thousands of weak clues? → yes → ML starts to earn its keepML wins when the pattern is too subtle, too wide, or too fast-moving for hand-written rules. Recognising faces, ranking search results, catching brand-new fraud styles — nobody can write those rules down.
Rules win when the pattern is known, stable, and small. Age limits, spending caps, working-hours checks, blocked-word lists.
A real example you have seen
Your bank blocks a card after three wrong PIN attempts. That is a rule, and it should stay a rule. The same bank uses machine learning to spot unusual spending, because "unusual" cannot be written down.
Remember this
- A rule is a fixed, human-written instruction. It is predictable, testable, and free to run.
- Use ML when the pattern is unknown, changes often, or hides in thousands of weak clues.
- Starting with a rule is not a failure. It is the baseline — the score every model must later beat.
What to learn next
- Turning a vague request into a prediction task — the next step once ML is genuinely warranted.
- The baselines you must beat first — the rule you almost wrote becomes the bar.
- What is machine learning? — if the sieve-versus-child distinction is still fuzzy.
Developer — Code and libraries.
Setup
pip install scikit-learnOutputs verified with scikit-learn 1.7 and numpy 1.26.
When the truth is a rule, the rule wins
Here the fraud pattern genuinely is "big money at odd hours". Watch a model try to rediscover it from data.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
rng = np.random.default_rng(42)
n = 2000
# Two facts the bank already stores about every transfer.
amount = rng.exponential(scale=3000, size=n).round(2)
hour = rng.integers(0, 24, size=n)
# The ops team's rule, written years ago: big money moving at odd hours.
def night_rule(amount, hour):
return ((amount > 9000) & ((hour <= 4) | (hour >= 23))).astype(int)
# In this world the rule IS the truth. Labels come straight from it.
labels = night_rule(amount, hour)
X = np.column_stack([amount, hour])
X_tr, X_te, y_tr, y_te = train_test_split(
X, labels, test_size=0.5, stratify=labels, random_state=0)
model = make_pipeline(StandardScaler(), LogisticRegression())
model.fit(X_tr, y_tr)
rule_acc = (night_rule(X_te[:, 0], X_te[:, 1]) == y_te).mean()
print(f"flagged transfers in test set: {y_te.sum()} of {len(y_te)}")
print(f"rule accuracy: {rule_acc:.3f}")
print(f"model accuracy: {model.score(X_te, y_te):.3f}")
caught = model.predict(X_te)[y_te == 1].sum()
print(f"frauds the model caught: {caught} of {y_te.sum()}")flagged transfers in test set: 14 of 1000 rule accuracy: 1.000 model accuracy: 0.984 frauds the model caught: 3 of 14
The walkthrough
The rule scores 1.000 because it is the process that made the labels. Real rules are rarely perfect, but a rule a domain expert trusts is often close.
The model scores 0.984 and looks fine — until you read the last line. It caught 3 frauds out of 14. Accuracy hides this because fraud is rare; a model that flags nothing at all would score 0.986 here. The lesson on model evaluation covers this trap in depth.
Why did the model struggle? The rule is an AND of two thresholds — a sharp rectangle in feature space. Logistic regression draws one straight line, so it cannot carve out a corner. A tree model would do better, and would still be second-best to the rule it is imitating.
The deeper point: the model needed 1,000 labelled examples to become a worse version of one sentence.
Common mistakes
Choosing ML because it is the exciting option. The interviewer asks about models; production rewards whatever works. Write the rule first, measure it, and let the model earn its place by beating it. The next section on baselines is built around this habit.
Choosing rules because ML feels risky, then drowning in rules. Twenty rules with exceptions that fight each other is how rule systems die. When you find yourself adding rule 21, that is the signal to revisit ML.
Forgetting the hybrid. Real systems layer both: hard rules for the known cases, a model for the grey zone. The rule "three wrong PINs blocks the card" runs before any model sees the data.
Testing the rule on data the rule created. If yesterday's rule decided which cases got investigated, your labels only exist where the rule fired. A model trained on those labels inherits the rule's blind spots.
Try it yourself
Change the labels so fraud depends on a pattern no one could write down: make it depend on (amount % 970) < 40 plus noise. Re-run both. The rule collapses and the model starts to earn its keep. Then ask: could anyone have known this without trying both?
What to learn next
- Turning a vague request into a prediction task — the next step once ML is genuinely warranted.
- The baselines you must beat first — the rule you almost wrote becomes the bar.
- What is machine learning? — if the sieve-versus-child distinction is still fuzzy.
Researcher — Mathematics and papers.
The decision, formalised
Choosing between a rule set and a learned model is a comparison of two function classes under a total cost:
$$ \text{cost}(f) = \mathbb{E}{(x,y) \sim \mathcal{D}}\left[\ell(f(x), y)\right] + C{\text{build}}(f) + C_{\text{maintain}}(f) $$
Where:
- $f$ — the deployed decision function, whether hand-written or learned.
- $\mathcal{D}$ — the data distribution, which drifts over time.
- $\ell$ — the task loss, in business units, not accuracy points.
- $C_{\text{build}}$ — labelling, training, and validation cost, roughly zero for a rule.
- $C_{\text{maintain}}$ — monitoring, retraining, and debugging cost. For ML this term dominates and is routinely underestimated.
A rule is a depth-limited decision list: high bias, near-zero variance, zero label cost. A learned model buys lower bias with every one of those costs. The scoping question is whether the bias reduction is worth the two $C$ terms.
Evidence that the maintenance term dominates
Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems (NeurIPS), is the standard reference. Its core claim: the model is a small box in a large system, and the system accumulates debt — entanglement, correction cascades, undeclared consumers, feedback loops — that rules mostly avoid.
Zinkevich's Rules of Machine Learning (Google, 2016) makes the ordering explicit. Rule 1 is "Don't be afraid to launch a product without machine learning." A heuristic gets you 50% of the benefit, defines the metric, and generates training data for the eventual model.
The selective-labels problem
When a deployed rule (or human judge) decides which cases receive outcomes, the observed labels are censored by the policy itself. Lakkaraju et al. (2017), The Selective Labels Problem (KDD), formalise this: offline evaluation of a replacement model on rule-filtered data is biased, and contraction techniques or randomised holdouts are needed. Any "replace the rule with a model" project inherits this problem on day one.
Where the frontier actually sits
- Interpretable-by-construction models — scoring systems (Ustun and Rudin, 2016, SLIM), falling-rule lists — occupy the middle ground: learned, yet auditable like rules.
- Rudin (2019), Stop Explaining Black Box Machine Learning Models for High Stakes Decisions (Nature Machine Intelligence), argues that for high-stakes tabular problems, the accuracy gap between interpretable models and black boxes is often negligible.
- Hybrid deployments (rules as guardrails around a model) are standard practice in fraud and credit; the rule layer bounds worst-case behaviour, which no validation-set metric can guarantee.
What to learn next
- Turning a vague request into a prediction task — the next step once ML is genuinely warranted.
- The baselines you must beat first — the rule you almost wrote becomes the bar.
- What is machine learning? — if the sieve-versus-child distinction is still fuzzy.