Testing ML Code and CI

Behavioural tests for models

A behavioural test checks what a trained model does on cases you choose on purpose, the way a driving examiner tests specific manoeuvres instead of only checking the car starts.

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. The three kinds
  5. How it works
  6. A real example you have seen
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A behavioural test checks specific things a trained model should and should not do, chosen on purpose, one at a time.

The analogy you have already lived

You have taken a driving test, or watched someone take one. The examiner does more than check the car starts and moves. They ask for a specific reverse parallel park, a specific hill start, a specific emergency stop. Each manoeuvre checks one skill, on purpose.

A single "did you pass, yes or no" score would hide which skill was missing. A behavioural test for a model works the same way: instead of one overall accuracy number, you check specific abilities, one at a time, by name.

Why it exists

A model can score 95% accuracy on a big test set. It can still fail in an obvious, embarrassing way on a case nobody happened to include in that set.

Averages hide this. A model can get every easy example right and every hard one wrong. It can still land on the same accuracy as a model that is uniformly mediocre everywhere. The single number cannot tell them apart, and only one of those two models is safe to ship.

Behavioural tests fix this by checking specific abilities directly, the way the driving examiner checks specific manoeuvres, instead of trusting one combined score to represent everything.

The three kinds

Minimum functionality — does the easiest possible version of the task work? "This is amazing" should be scored positive. If this fails, nothing else about the model matters yet.

Invariance — does something the model should ignore actually get ignored? Swapping "restaurant" for "movie" in "I loved this ___" should not change whether the sentence is positive.

Directional expectation — does a change with an obvious correct direction move the output the right way? Adding the word "not" in front of a positive word should make the sentence read as negative.

How it works

   trained model
        |
        |--- minimum functionality test  ---> "this is amazing"      should be POSITIVE
        |--- invariance test              ---> swap an unrelated word, label should NOT change
        |--- directional expectation test ---> add "not", label SHOULD change
        |
        v
   three separate answers, not one blended score

A real example you have seen

A translation app that handles "I am happy" correctly but mistranslates "I am not happy" as still positive has failed a directional expectation test. Many real translation and sentiment systems have failed this exact test in public, because nobody checked that specific ability on purpose.

The honest part

Some behavioural failures are genuinely hard to fix. A model can fail an invariance or directional test not because of a silly bug, but because of a real limit in what it learned. The developer section below shows exactly that happening, and shows why the quick fix does not work either. Read it in full before assuming any behavioural bug has an easy patch.

Remember this

  • Behavioural tests check specific abilities, chosen on purpose, instead of one blended accuracy number.
  • The three kinds are minimum functionality, invariance, and directional expectation.
  • A model can pass on accuracy and still fail an obvious behavioural test — that gap is exactly what these tests are for.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn joblib pytest

The model being tested

A tiny sentiment classifier: TF-IDF turns each sentence into word-count features, and logistic regression learns which words push toward "positive" or "negative". Twenty short sentences, ten of each label.

sentiment.py
"""A tiny sentiment classifier, small enough to test its behaviour directly."""
import joblib
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline

TRAIN_TEXTS = [
    "amazing product, loved it", "terrible experience, would not buy again",
    "great quality and fast delivery", "awful, broke on day one",
    "works perfectly, very happy", "waste of money, do not buy",
    "excellent value for the price", "disappointing and slow",
    "fantastic support team", "poor build quality",
    "highly recommend this", "would not recommend at all",
    "best purchase this year", "worst purchase this year",
    "solid and reliable", "flimsy and cheap feeling",
    "happy with the results", "unhappy with the results",
    "good for the price", "bad for the price",
]
TRAIN_LABELS = [1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 0]

model = make_pipeline(TfidfVectorizer(), LogisticRegression(max_iter=1000))


def train():
    model.fit(TRAIN_TEXTS, TRAIN_LABELS)
    joblib.dump(model, "sentiment.joblib")
    return model


if __name__ == "__main__":
    m = train()
    print("train accuracy:", round(m.score(TRAIN_TEXTS, TRAIN_LABELS), 4))
bash
python sentiment.py
Output
train accuracy: 1.0

Perfect accuracy on the training sentences is expected here — with only twenty short, plainly-worded sentences and no noise, the model can memorise the pattern. It says nothing about whether the model behaves sensibly on sentences it has never seen. That is what the tests below check.

The behavioural tests

test_behaviour.py
import joblib
import pytest

MODEL = joblib.load("sentiment.joblib")


def predict(text: str) -> int:
    return int(MODEL.predict([text])[0])


# --- Minimum functionality: the easiest possible cases must work ---

@pytest.mark.parametrize("text,expected", [
    ("this is amazing", 1),
    ("this is terrible", 0),
    ("absolutely fantastic", 1),
    ("absolutely awful", 0),
])
def test_minimum_functionality(text, expected):
    assert predict(text) == expected


# --- Invariance: swapping an unrelated noun must not change the label ---

def test_invariant_to_the_thing_being_praised():
    assert predict("I loved this restaurant") == predict("I loved this movie")


# --- Directional expectation: negating a word should flip the label ---

def test_negation_flips_a_positive_word():
    assert predict("this is not good") != predict("this is good")
bash
pytest test_behaviour.py -q
Output
.....F                                                                   [100%]
================================== FAILURES ===================================
_____________________ test_negation_flips_a_positive_word _____________________

    def test_negation_flips_a_positive_word():
>       assert predict("this is not good") != predict("this is good")
E       AssertionError: assert 1 != 1
E        +  where 1 = predict('this is not good')
E        +  and   1 = predict('this is good')

test_behaviour.py:32: AssertionError
=========================== short test summary info ===========================
FAILED test_behaviour.py::test_negation_flips_a_positive_word - AssertionErro...
1 failed, 5 passed in 1.71s

That is a real run against a real model. Four minimum-functionality cases pass. The invariance check passes: swapping "restaurant" for "movie" does not move the score. The negation test genuinely fails — the model calls "this is not good" positive, same as "this is good".

Why this failure is real, not a bug you can quickly patch

TF-IDF counts words. It has no idea that "not" sitting in front of "good" changes what the sentence means — to a word-counting model, "not good" looks mostly like the word "good" plus one common, low-signal word, "not". The model was never shown a single negated sentence during training, so it never had a chance to learn that pattern.

The obvious next step is adding a few negated examples and retraining. Tried here, honestly, with the actual result:

python
extra_texts = ["this is not good", "this is not amazing", "not a good product", "this is not bad", "this is not terrible"]
extra_labels = [0, 0, 0, 1, 1]
# retrain on TRAIN_TEXTS + extra_texts, TRAIN_LABELS + extra_labels
Output
this is good      -> 0  (was 1 before adding the examples)
this is not good  -> 0
this is bad       -> 1  (was 0 before adding the examples)
this is not bad   -> 0

Adding five examples to a twenty-sentence dataset did not teach negation — it moved the decision boundary somewhere else confusing instead, and now plain "this is good" is called negative too. This part is genuinely hard, and there is no quick fix here — with this little data and a model that only counts words, negation is not learnable. A real fix needs either a much larger, negation-rich training set, or a model that reads words in order rather than only counting them, such as a fine-tuned BERT classifier.

Common mistakes

Writing behavioural tests only for cases the model already handles well. The value is in testing the cases you are unsure about, particularly the ones a plain accuracy score would never surface.

Treating a failed behavioural test as a bug ticket by default. Sometimes it reveals a real architectural limit, as above. The right response is a documented, known limitation — not a rushed patch that quietly breaks something else, the way the negation retrain did here.

Only testing text models this way. The same three kinds apply to tabular models — CI/CD for machine learning shows a directional expectation test on a loan-approval model, where raising income should never lower the approval score.

Confusing invariance with "the model ignores everything." An invariance test checks one specific, chosen irrelevant change. It is not a claim that the model should ignore all variation — only the variation you deliberately picked because it should not matter.

Try it yourself

Add an invariance test that swaps a name: predict("I loved this from Rahul") versus predict("I loved this from Priya"). Run it. If it fails, you have found a real bias in this small model worth writing down — not fixing blindly, writing down, so whoever ships it knows about it.

What to learn next

Researcher — Mathematics and papers.

The framework this lesson is built on

Ribeiro et al. (2020), Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, introduces exactly the three test types used above — MFT (minimum functionality), INV (invariance), and DIR (directional expectation) — plus a taxonomy of linguistic capabilities to probe: vocabulary, named entity recognition, negation, coreference, semantic role labelling, and robustness to typos. Their user study found that engineers using CheckList on a mature, well-evaluated commercial sentiment model discovered roughly five times as many distinct bugs as the same engineers found using only held-out test-set accuracy.

Negation is explicitly called out in that paper as one of the capabilities most bag-of-words and even early neural models fail, for exactly the mechanistic reason demonstrated above: a model with no notion of word order or scope cannot represent that "not" inverts the polarity of what follows it.

Metamorphic relations as the general form

Chen et al.'s (1998) metamorphic testing is the general concept behind INV and DIR: rather than asserting a specific output for an input (which requires a ground-truth label), you assert a relation between the outputs of two related inputs. This sidesteps the oracle problem — you do not need to know the correct probability for "this is not good," only that it should differ from the probability for "this is good."

For tabular models the same idea produces monotonicity tests (income up, score never down), permutation tests (row order should not matter), and additivity tests for linear-in-the-features models.

Where behavioural failure sits relative to accuracy

Held-out accuracy is a marginal statistic — an average over the test distribution. A capability failure that affects 3% of real traffic can be invisible in an accuracy number computed on a test set where that pattern appears rarely, while being catastrophic for the users who hit it. This is the formal reason averages hide the failures behavioural tests are built to find: a low-frequency subgroup failure moves an aggregate metric by less than its own sampling noise.

Papers

  • Ribeiro et al., Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, ACL 2020 — arxiv.org/abs/2005.04118
  • Chen, Cheung and Yiu, Metamorphic Testing: A New Approach for Generating Next Test Cases, 1998
  • Naik et al., Stress Test Evaluation for Natural Language Inference, COLING 2018 — an earlier, narrower precursor covering negation and word-overlap stress tests specifically.

What to learn next