Releasing Models Safely

Rolling back a model

Rolling back means switching straight back to the last version you trusted, without stopping to diagnose the new one first — the same instinct as putting your old shoes back on the moment new ones start to hurt.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Rolling back means switching straight back to the last version you trusted, before you stop to figure out what went wrong.

The analogy you have already lived

You have bought a new pair of shoes, worn them out once, and felt them start to hurt. You did not stand in the street trying to work out exactly which stitch was wrong. You went home and put your old shoes back on. Working out what was wrong with the new pair came later, at your own pace, with your feet no longer hurting.

Rolling back a model is that same instinct. The moment something is unmistakably wrong, switch back to the version that was working — first. Investigate why, after.

Why it exists

The natural reaction to a broken release is to try to fix it, live, immediately. That instinct is usually wrong. Diagnosing a real problem under pressure, while it is actively affecting users, is slow and error-prone. Those are exactly the conditions under which people make a second mistake on top of the first one.

Rolling back removes the time pressure from the diagnosis. The moment you switch back to a version you trust, the bleeding stops. Nobody is being affected any more while you work out, calmly, what actually went wrong.

How it works

   new model deployed
          |
          v
   something looks wrong -- errors up, a metric moved, a support ticket pattern
          |
          v
   ROLL BACK FIRST                 (switch straight back to the last trusted version)
          |
          v
   users are on a version you trust again, right now
          |
          v
   THEN investigate                (with no pressure, no users currently affected)

A real example you have seen

An app update that turns out to be badly broken usually gets pulled from the app store and replaced with the previous version within hours. This happens long before the developers have fully worked out what caused the crash. The rollback and the root-cause investigation are two separate steps, done in that order, not one.

The honest part

Rolling back only works if there is genuinely something safe to roll back to, and a fast way to get there. Both of those have to be built and kept ready before you need them — not invented during an incident.

Remember this

  • Rolling back means reverting first, investigating after — not the other way around.
  • It only works if a known-good previous version is kept ready, on purpose, in advance.
  • A rollback should go to whichever version was active right before this one, not always to the very first release.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pytest

No external libraries needed — the core idea is a small, careful piece of bookkeeping, not a specific tool.

Keeping a real release history

The detail that matters most: rollback must go to whichever version was active immediately before the current one — not to the first version ever shipped. A history that only remembers "the current one" and "version one" gets this wrong the second time a rollback is ever needed.

release_history.py
"""Keeps a small, ordered history of releases so a rollback always goes to
the version that was actually active immediately before the current one --
not to the first version ever shipped.
"""
import time


class ReleaseHistory:
    def __init__(self):
        self.deployed_order = []   # every version, in the order it was deployed
        self.active_version = None
        self.log = []

    def deploy(self, version: str):
        self.deployed_order.append(version)
        previous = self.active_version
        self.active_version = version
        self.log.append({
            "time": round(time.time(), 3), "action": "deploy",
            "from": previous, "to": version, "reason": "manual release",
        })

    def rollback(self, reason: str):
        """Goes back to whichever version was active immediately before this one."""
        if len(self.deployed_order) < 2:
            raise RuntimeError("no previous version to roll back to")
        current = self.active_version
        previous = self.deployed_order[self.deployed_order.index(current) - 1]
        self.active_version = previous
        self.log.append({
            "time": round(time.time(), 3), "action": "rollback",
            "from": current, "to": previous, "reason": reason,
        })
        return previous


def monitor_and_maybe_rollback(history: ReleaseHistory, error_rate: float, threshold: float = 0.05):
    if error_rate > threshold:
        new_active = history.rollback(reason=f"error rate {error_rate:.3f} exceeded threshold {threshold}")
        print(f"AUTO-ROLLBACK: {history.log[-1]['from']} -> {new_active}  ({history.log[-1]['reason']})")
        return True
    print(f"healthy: error rate {error_rate:.3f} is within the {threshold} threshold, no action")
    return False

Two releases, two rollbacks, run for real

python
from release_history import *

history = ReleaseHistory()
history.deploy("v1")
history.deploy("v2")

# v2 turns out to have a real problem
monitor_and_maybe_rollback(history, error_rate=0.11, threshold=0.05)
print("active after monitoring:", history.active_version)

# later, v3 ships (a fixed version) and later ALSO needs a rollback --
# it must go back to v2, not all the way back to v1
history.deploy("v3")
monitor_and_maybe_rollback(history, error_rate=0.09, threshold=0.05)
print("active after second rollback:", history.active_version)
Output
AUTO-ROLLBACK: v2 -> v1  (error rate 0.110 exceeded threshold 0.05)
active after monitoring: v1
AUTO-ROLLBACK: v3 -> v2  (error rate 0.090 exceeded threshold 0.05)
active after second rollback: v2

The second rollback correctly lands on v2, not v1 — even though v1 was the very first version ever deployed. A history that only tracked "the current version" and "the original version" would have gotten this wrong, silently reverting further than intended.

The full log, for a real postmortem

python
from release_history import *

history = ReleaseHistory()
history.deploy("v1")
history.deploy("v2")
history.rollback(reason="error rate 0.110 exceeded threshold 0.05")
history.deploy("v3")
history.rollback(reason="error rate 0.090 exceeded threshold 0.05")

for entry in history.log:
    print(entry)
Output
{'time': 1788334978.349, 'action': 'deploy', 'from': None, 'to': 'v1', 'reason': 'manual release'}
{'time': 1788334978.349, 'action': 'deploy', 'from': 'v1', 'to': 'v2', 'reason': 'manual release'}
{'time': 1788334978.349, 'action': 'rollback', 'from': 'v2', 'to': 'v1', 'reason': 'error rate 0.110 exceeded threshold 0.05'}
{'time': 1788334978.349, 'action': 'deploy', 'from': 'v1', 'to': 'v3', 'reason': 'manual release'}
{'time': 1788334978.349, 'action': 'rollback', 'from': 'v3', 'to': 'v2', 'reason': 'error rate 0.090 exceeded threshold 0.05'}

Every time value here is a real time.time() timestamp from the machine that ran this — yours will show today's date instead, and the five entries will not be identical to a fraction of a second apart on a slower machine. The fields that matter for a postmortem — action, from, to, reason — will match exactly.

Every rollback records exactly what it reverted from, what it reverted to, and why — the exact three facts a postmortem needs, captured automatically at the moment of the incident rather than reconstructed from memory afterward.

Testing the rollback logic

test_rollback.py
import pytest
from release_history import ReleaseHistory, monitor_and_maybe_rollback


def test_rollback_goes_to_the_immediately_prior_version_not_the_first_one():
    history = ReleaseHistory()
    history.deploy("v1")
    history.deploy("v2")
    history.deploy("v3")
    history.rollback(reason="test")
    assert history.active_version == "v2"


def test_rollback_with_no_previous_version_raises_instead_of_silently_doing_nothing():
    history = ReleaseHistory()
    history.deploy("v1")
    with pytest.raises(RuntimeError):
        history.rollback(reason="test")


def test_a_bad_error_rate_triggers_an_automatic_rollback():
    history = ReleaseHistory()
    history.deploy("v1")
    history.deploy("v2")
    triggered = monitor_and_maybe_rollback(history, error_rate=0.20, threshold=0.05)
    assert triggered is True
    assert history.active_version == "v1"


def test_every_rollback_is_recorded_with_a_reason():
    history = ReleaseHistory()
    history.deploy("v1")
    history.deploy("v2")
    history.rollback(reason="manual: bad answers reported by support")
    last = history.log[-1]
    assert last["action"] == "rollback"
    assert "bad answers" in last["reason"]
bash
pytest test_rollback.py -q
Output
.....                                                                    [100%]
5 passed in 0.02s

Common mistakes

Not actually keeping the previous artifact. A rollback plan is worthless if the previous model file was deleted, overwritten, or never saved anywhere. Keep at least the last few versions available, exactly as covered in model registries.

Diagnosing before reverting. The instinct to understand a problem before acting feels responsible, and it is usually the wrong order under real pressure. Revert to safety first; the investigation does not need users to still be affected while it happens.

Rolling back the model but not the data it already changed. If the bad version wrote predictions, logs, or state that something downstream already consumed, switching the model back does not undo that. This is the "one-way door" problem — some effects a rollback genuinely cannot reverse, and it is worth knowing in advance which of your effects are one-way doors and which are not.

No automatic trigger, only a manual one. A human has to notice something is wrong before they can act on it, and noticing takes time. Pairing rollback with an automatic trigger — as monitor_and_maybe_rollback does above — removes that human reaction-time cost from the most urgent cases, while still allowing a manual rollback for anything the automatic check does not catch.

Assuming "rollback" always means "go back one step". As the second scenario above shows, "one step back" from v3 is v2, not v1. A shallow implementation that only remembers the very first version gets this wrong exactly when it matters most — the second time.

Try it yourself

Add a max_rollback_depth check: if the same version was already rolled back to within the last hour, refuse a second automatic rollback and instead raise a loud alert for a human. This catches a real failure mode — a model that gets deployed, rolled back, redeployed unchanged, and rolled back again, in a loop nobody is actually looking at.

What to learn next

Researcher — Mathematics and papers.

Rollback as a compensating action, not a true undo

In distributed-systems terms, a model rollback is a compensating transaction: it does not undo the effects of the bad version, it applies a new action (switch back) whose net effect approximates reversal for the specific state the rollback controls — the active model pointer. Anything outside that state (writes to a database, messages published to a downstream queue, cached results computed under the bad model) is untouched by the rollback and may require its own, separate compensating action, or may be irreversible entirely. Identifying which effects of a release are reversible by a rollback and which are not is a design question that belongs in the release plan, not something to discover mid-incident.

Rollback safety and the two-version contract

A rollback is only safe if the previous version can correctly handle whatever state the current version left behind — most concretely, whatever the current version wrote to a shared store. This is the same backward compatibility constraint discussed for schema and API changes in versioning a model API: a rollback plan implicitly assumes N and N-1 can coexist, or at minimum that N-1 can safely read whatever N most recently wrote. Deploying a version whose data writes an older version cannot understand removes the rollback option in practice, even if the model artifact itself is readily available.

Automatic rollback as a control system

Framed as a feedback controller, monitor_and_maybe_rollback implements a simple bang-bang controller: below a threshold, no action; above it, an immediate, maximal corrective action (full revert). This is deliberately simpler than a gradual controller (like a canary's step-wise ramp) because the situation it responds to — an active incident — calls for the fastest possible return to a known-safe state, not a careful, incremental correction. The trade-off is sensitivity to noisy metrics: a threshold set too tight on a noisy error-rate signal triggers rollbacks on statistical noise rather than real regressions, the same false-positive concern raised for evaluation gates and canary guardrails.

Mean time to recovery as the metric that matters

Site-reliability practice (Beyer et al., Site Reliability Engineering, 2016) treats MTTR (mean time to recovery) as often more consequential to user-facing impact than mean time between failures — a team that rolls back in ninety seconds causes far less harm per incident than one that takes forty-five minutes to diagnose forward, even if the second team has fewer incidents overall. This is the operational argument underneath "revert first, diagnose after": it directly minimises the metric that determines how much an incident actually costs.

Papers and further reading

  • Beyer, Jones, Petoff and Murphy (eds.), Site Reliability Engineering, O'Reilly 2016 — the MTTR framing and the general case for automated, fast rollback over careful live diagnosis.
  • Fowler, CompensatingTransaction — the distributed-systems framing of a rollback as an approximate, not exact, undo.

What to learn next

What to learn next

These follow on from what you just read.

  • Releasing Models Safely

    Feature flags for model rollouts

    A feature flag turns one specific behaviour on or off by editing a config value, the way a household electrical panel lets you cut power to only the kitchen without touching any other circuit — or the main switch.

  • Releasing Models Safely

    Champion and challenger models

    A champion model keeps serving every real decision while one or more challengers are scored continuously alongside it, and only take over after a sustained, statistically real advantage — the same discipline as a table-tennis champion who keeps the table only while they keep winning, not after one lucky point.

  • Releasing Models Safely

    Swapping weights with no downtime

    A zero-downtime weight swap replaces a model's numbers inside an already-running server, the way a relay runner hands off the baton at full speed — the new runner is carrying it before the old one lets go, so the race never actually stops.