Releasing Models Safely

Feature flags for model rollouts

A feature flag turns one specific behaviour on or off by editing a config value, the way a household electrical panel lets you cut power to only the kitchen without touching any other circuit — or the main switch.

On this page 7
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A feature flag turns one specific behaviour on or off by changing a setting, without deploying new code.

The analogy you have already lived

You have stood in front of a household electrical distribution board — the panel of small switches, each one controlling a single circuit. Flip one, and only the kitchen loses power; the lights, the fridge, everything else stays on. You do not need to touch the main switch, and you do not need an electrician to rewire the house to do it.

A feature flag is that same panel, for software. "Use the new model" can be its own switch, flipped on for a few people, or everyone, or nobody. It works independently of every other part of the system, and independently of whether new code was recently deployed.

Why it exists

Rolling back a model covered reverting an entire release. That works, but it is a blunt instrument. It reverts everything the release contained, all at once, and doing it usually still means a deploy action of some kind.

A feature flag separates two things that are easy to accidentally treat as one: shipping code, and turning a behaviour on. New code, including a new model's calling logic, can be deployed while switched off — sitting there, doing nothing, completely safe. Turning it on later is a config change, not a new deployment. That makes it something you can do in seconds, and undo in seconds, independent of your entire release pipeline.

How it works

   deploy the new code, flag OFF
        |
        v
   code is live, but is not being used by anyone yet -- this is "dark"
        |
        v
   flip the flag on for a small allowlist (your own team, first)
        |
        v
   flip the flag to a rollout percentage: 10% -> 50% -> 100%
        |
        v
   something wrong at any point?  flip the flag back off.
   instantly. no redeploy. no rollback of any code at all.

A real example you have seen

Apps often show a new feature to some users and not others — a redesigned checkout page, a new recommendation panel. They do this without releasing a new version of the app itself. The code for the new feature was already sitting on your phone. A setting on the company's server decided, on that day, whether it was switched on for you.

Remember this

  • A feature flag separates shipping code from turning it on — two different actions, not one.
  • Flipping a flag is a config change, which is why it can be instant, in both directions.
  • It gives you a lever finer than "this whole release" — a single behaviour, for a chosen slice of users, independent of everything else.

What to learn next

  • Champion and challenger models — running two models side by side over a longer window than a typical flag rollout, to decide which one wins.
  • Circuit breakers and fallbacks — a flag-adjacent pattern for automatically falling back when a model call itself is failing.
  • Versioning a model API — the complementary discipline for changes a flag alone cannot safely gate, like a breaking change to the request or response shape.

Developer — Code and libraries.

Setup

bash
pip install pytest

No external service needed for the idea itself — a real flag system (LaunchDarkly, Unleash, a home-grown config table) adds caching, audit logs, and a UI on top of exactly this mechanism.

The flag system

Reuses the same sticky, hash-based percentage rollout from canary releases — the mechanism is genuinely the same idea, applied at a finer grain than an entire model version.

flags.py
"""A small feature-flag system: turns a behaviour on or off, for specific
users or a percentage of traffic, by editing a config file -- never by
deploying new code.
"""
import hashlib
import json
import os

FLAGS_PATH = "flags.json"


def load_flags() -> dict:
    if not os.path.exists(FLAGS_PATH):
        return {}
    with open(FLAGS_PATH) as f:
        return json.load(f)


def save_flags(flags: dict):
    with open(FLAGS_PATH, "w") as f:
        json.dump(flags, f, indent=2)


def is_enabled(flag_name: str, user_id: str) -> bool:
    flags = load_flags()
    rule = flags.get(flag_name)
    if rule is None or not rule.get("enabled", False):
        return False
    if user_id in rule.get("allowlist", []):
        return True
    percent = rule.get("rollout_percent", 0)
    digest = hashlib.md5(f"{flag_name}:{user_id}".encode()).hexdigest()
    bucket = int(digest[:8], 16) % 10_000
    return bucket < percent * 100

Ramping a rollout, purely by editing a config value

python
from flags import *

save_flags({"use_new_model": {"enabled": True, "rollout_percent": 0, "allowlist": ["internal-tester-1"]}})

users = [f"user-{i}" for i in range(20_000)]

for target_pct in [0, 10, 50, 100]:
    flags = load_flags()
    flags["use_new_model"]["rollout_percent"] = target_pct
    save_flags(flags)
    on = sum(is_enabled("use_new_model", u) for u in users)
    print(f"rollout_percent={target_pct:>3}  ->  measured on for {on}/{len(users)} users ({on/len(users)*100:.2f}%)")
Output
rollout_percent=  0  ->  measured on for 0/20000 users (0.00%)
rollout_percent= 10  ->  measured on for 2025/20000 users (10.12%)
rollout_percent= 50  ->  measured on for 10000/20000 users (50.00%)
rollout_percent=100  ->  measured on for 20000/20000 users (100.00%)

Not one line of application code changed between any of these four steps — every step is the exact same save_flags config edit, at a different percentage.

The kill switch, timed for real

python
from flags import *
import time

t0 = time.perf_counter()
flags = load_flags()
flags["use_new_model"]["enabled"] = False
save_flags(flags)
print(f"flag flip took: {(time.perf_counter()-t0)*1000:.3f} ms")
print("random-user-77 after kill switch:", is_enabled("use_new_model", "random-user-77"))
Output
flag flip took: 0.372 ms
random-user-77 after kill switch: False

That 0.372 ms is one real, tiny file write on one machine — a real production flag system typically reads from a fast cache rather than a file on every call, so the number itself is not the point. What matters is the kind of operation: a config write, not a deployment. For comparison, redeploying even a modest fleet of forty warm serving instances at roughly fifty milliseconds of model-load time each — the figure measured for a single load in blue-green model deployments — works out to about two full seconds done one at a time, and that figure ignores the deploy pipeline itself, which is usually the larger cost. The exact multiplier will differ for your infrastructure; the gap in kind between a config write and a deploy will not.

Testing the flag logic

test_flags.py
from flags import is_enabled, save_flags


def test_a_disabled_flag_is_off_for_everyone_even_the_allowlist():
    save_flags({"f": {"enabled": False, "rollout_percent": 100, "allowlist": ["vip"]}})
    assert is_enabled("f", "vip") is False


def test_allowlisted_users_are_on_regardless_of_percent():
    save_flags({"f": {"enabled": True, "rollout_percent": 0, "allowlist": ["vip"]}})
    assert is_enabled("f", "vip") is True


def test_a_100_percent_rollout_is_on_for_everyone():
    save_flags({"f": {"enabled": True, "rollout_percent": 100, "allowlist": []}})
    assert all(is_enabled("f", f"user-{i}") for i in range(200))


def test_an_unknown_flag_defaults_to_off():
    save_flags({})
    assert is_enabled("does_not_exist", "anyone") is False
bash
pytest test_flags.py -q
Output
....                                                                     [100%]
4 passed in 0.06s

test_a_disabled_flag_is_off_for_everyone_even_the_allowlist is worth reading twice — it checks the master enabled switch overrides even a hand-picked allowlist. This is the actual kill switch: one field, checked first, before anything else about the flag is even considered.

Common mistakes

Deploying and flipping the flag on in the same action. That throws away the entire benefit — the point of a flag is to separate "the code exists and is safe" from "the code is active", and doing both at once collapses that separation back into an ordinary all-at-once release.

Forgetting a flag exists, months later. A flag left at 100% forever, with the old code path never removed, is a form of debt: two code paths to maintain, one of which nobody is testing any more because it is "always on". Retire a flag — remove the old path entirely — once a rollout is complete and stable.

No default when the flag system itself fails. is_enabled above returns False for a flag it cannot find, on purpose. A flag check should always fail toward the safer, more conservative behaviour, not silently enable something because a lookup errored.

Using a flag as a substitute for real access control. A flag decides whether a feature is active; it is not designed or audited as a security boundary. Do not rely on a flag alone to hide something a determined user genuinely should not be able to reach.

Too many flags interacting unpredictably. Five flags, each independently on or off, is thirty-two possible combinations of active behaviour — most of which nobody has ever actually tested. Keep the number of flags active at once small, and remove them promptly once their purpose is served.

Try it yourself

Add a second flag, use_new_model_batch_endpoint, and write a test proving the two flags can be toggled completely independently — one on, one off, in either combination — without any code change. This is the property that makes flags genuinely finer-grained than a whole-model rollback.

What to learn next

  • Champion and challenger models — running two models side by side over a longer window than a typical flag rollout, to decide which one wins.
  • Circuit breakers and fallbacks — a flag-adjacent pattern for automatically falling back when a model call itself is failing.
  • Versioning a model API — the complementary discipline for changes a flag alone cannot safely gate, like a breaking change to the request or response shape.

Researcher — Mathematics and papers.

Flags as a form of runtime configuration, not deployment

Formally, a feature flag moves a decision from compile/deploy time to runtime, evaluated on every request against externally mutable state. This is a specific application of the broader configuration-over-code principle: any decision that might need to change faster than your deploy cycle allows is a candidate for extraction into a flag, provided the check itself is cheap enough to evaluate on the request path without becoming its own latency problem — the hashing approach shared with canary routing keeps this check to microseconds.

Flag debt and combinatorial testing cost

Every active flag doubles the number of distinct code paths a system can be in. With $n$ independent flags, there are $2^n$ possible combined states, and most real test suites exercise a small fraction of them — usually only "all off" and "all on". This is the formal version of the "too many flags" mistake above: the untested combinations are not hypothetical, they are the states production will eventually visit as flags are toggled independently over time, and the ones a test suite has never seen are exactly where a real incident tends to originate.

Flags versus traffic-splitting infrastructure

A percentage-based flag and a canary router (from canary releases for models) implement the same hashing mechanism, and the distinction between them is organisational rather than technical: a canary typically governs an entire service version, evaluated once at the load-balancer or router layer, while a flag typically governs one narrow behaviour inside a single running version, evaluated per-decision inside the application. Large-scale flag platforms (LaunchDarkly, Unleash, homegrown systems at bigger companies) converge on the same primitive — a targeting rule evaluated against a stable identifier — because the underlying problem, consistent per-entity assignment under a changeable rule, is identical in both cases.

Flags in a machine-learning-specific context

Beyond gating "use the new model", flags are commonly used for narrower ML-specific toggles: which prompt template an LLM feature uses, whether a fallback heuristic activates when a model call fails or times out, or which of several candidate feature sets a scoring pipeline reads — see circuit breakers and fallbacks for the fallback case specifically. In each case the value of the flag is the same: the behaviour can be changed, or reverted, without waiting on the model-training or deployment pipeline that produced the artifact it is gating.

Papers and prior art

  • Fowler, FeatureToggle, martinfowler.com — martinfowler.com/bliki/FeatureToggle.html — the canonical taxonomy (release toggles, experiment toggles, ops toggles, permission toggles) that this lesson's use_new_model flag sits under as a release toggle.
  • Rahman et al., Feature Toggles: Practitioner Practices and a Case Study, MSR 2016 — an empirical study of flag debt and toggle lifecycle in real codebases.

What to learn next

  • Champion and challenger models — running two models side by side over a longer window than a typical flag rollout, to decide which one wins.
  • Circuit breakers and fallbacks — a flag-adjacent pattern for automatically falling back when a model call itself is failing.
  • Versioning a model API — the complementary discipline for changes a flag alone cannot safely gate, like a breaking change to the request or response shape.

What to learn next

These follow on from what you just read.

  • Releasing Models Safely

    Champion and challenger models

    A champion model keeps serving every real decision while one or more challengers are scored continuously alongside it, and only take over after a sustained, statistically real advantage — the same discipline as a table-tennis champion who keeps the table only while they keep winning, not after one lucky point.

  • Releasing Models Safely

    Swapping weights with no downtime

    A zero-downtime weight swap replaces a model's numbers inside an already-running server, the way a relay runner hands off the baton at full speed — the new runner is carrying it before the old one lets go, so the race never actually stops.

  • Releasing Models Safely

    Versioning a model API

    Versioning a model API means freezing the exact shape of a contract once real callers depend on it, the way you keep an old phone number forwarding for a while after getting a new one, so nobody who still has it suddenly cannot reach you.