Scaling and Traffic Management

Graceful shutdown and connection draining

A server that is about to stop should finish the requests it already started, and refuse new ones, instead of cutting everyone off at once.

Read these first

On this page 9
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. Two things a graceful shutdown does, in order
  5. How it works
  6. A real example you have seen
  7. The honest part
  8. Remember this
  9. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Graceful shutdown means a server stops accepting new requests but finishes the ones it already started, before it actually exits.

The analogy you have already lived

You have been inside a shop at closing time. The staff do not switch off the lights and walk out on the customers already at the till. They lock the door, put up a sign, and finish serving whoever is already inside.

New customers see the sign and know to come back tomorrow. The ones already inside get served properly, without being cut off mid-purchase.

Why it exists

Servers get shut down constantly, on purpose. A new version is deploying, a machine is scaling down, a crashed process is being replaced. Every one of those is a normal, routine event, not an emergency.

If a server exits the instant it is told to stop, every request it was handling gets abandoned. The caller sees a broken connection, not an answer. That request may well have been about to succeed.

Two things a graceful shutdown does, in order

Stop accepting new requests. The server marks itself as "not accepting," so nothing new starts arriving. A load balancer or router in front of it should already have stopped sending traffic here.

Wait for the requests already in progress to finish, called draining — then, and only then, actually exit.

How it works

   shutdown signal received
        |
        v
   [ stop accepting new requests ]   <- new requests get rejected immediately
        |
        v
   [ wait for in-flight requests to finish ]   <- draining
        |
        v
   [ exit, safely ]

A server with no drain step stops mid-request, mid-response, with no warning to whoever was talking to it.

A real example you have seen

Almost any app update on your phone happens without you noticing a request fail. The old version finishes serving whatever it was doing. The new version takes over new requests, and the switch is invisible. That invisibility is graceful shutdown working exactly as intended.

The honest part

Draining has to have a limit. A request that hangs forever would keep the old server alive forever too. That blocks the very deployment that triggered the shutdown. Every real system pairs draining with a maximum wait — after which it exits anyway, incomplete requests and all.

Remember this

  • Graceful shutdown has two steps: stop accepting new requests, then finish the ones already running.
  • Draining is the waiting period between those two steps.
  • Draining always needs a maximum wait, or one stuck request can block a shutdown indefinitely.

What to learn next

Developer — Code and libraries.

Setup

No installs needed beyond the standard library.

Watching a real drain happen, in real time

This runs actual threads with actual timing, not a simulation. Draining is fundamentally about real concurrent behaviour — what is genuinely in-flight at the moment the shutdown signal lands.

graceful_shutdown.py
import threading
import time

SERVICE_TIME = 0.3   # how long one request takes to handle
DRAIN_AT = 1.0        # shutdown signal arrives at this simulated moment
NUM_REQUESTS = 30
GAP = 0.05             # requests arrive 50ms apart

draining = False
in_flight = 0
lock = threading.Lock()
completed = []
rejected = []


def handle_request(req_id, arrive_at):
    global in_flight
    time.sleep(max(0, arrive_at - (time.perf_counter() - start_time)))
    with lock:
        if draining:
            rejected.append(req_id)
            return
        in_flight += 1
    time.sleep(SERVICE_TIME)  # doing the actual work
    with lock:
        in_flight -= 1
        completed.append(req_id)


def begin_drain():
    global draining
    time.sleep(DRAIN_AT)
    draining = True
    print(f"[drain] shutdown signal received, {in_flight} request(s) already in flight")
    while True:  # a real server waits here until in_flight reaches 0 before exiting
        with lock:
            if in_flight == 0:
                break
        time.sleep(0.01)
    print("[drain] all in-flight requests finished, safe to exit")


start_time = time.perf_counter()
drain_thread = threading.Thread(target=begin_drain)
drain_thread.start()

workers = [threading.Thread(target=handle_request, args=(i, i * GAP)) for i in range(NUM_REQUESTS)]
for w in workers:
    w.start()
for w in workers:
    w.join()
drain_thread.join()

print(f"\ncompleted: {len(completed)}")
print(f"rejected after drain started: {len(rejected)}")
print(f"total: {len(completed) + len(rejected)} of {NUM_REQUESTS}")
Output
[drain] shutdown signal received, 7 request(s) already in flight
[drain] all in-flight requests finished, safe to exit

completed: 21
rejected after drain started: 9
total: 30 of 30

This uses real threads and real wall-clock timing, so the exact counts can shift by one request on your machine. A request landing right at the t=1.0s boundary can go either way, depending on thread scheduling. Running this three times on the development machine produced 21/9, 21/9, and 20/10 completed/rejected. The pattern is the real point: some requests finish, later ones are rejected, none are silently dropped mid-request. It held every time.

Walking through it

The draining check happens before in_flight is incremented. A request that arrives after the signal never gets counted as in-flight at all. It is rejected cleanly, before any work starts on it.

begin_drain blocks on in_flight == 0 before printing that it is safe to exit. This is the actual draining step: the shutdown does not proceed until every request that was already running has genuinely finished.

None of the 30 requests are unaccounted for. completed + rejected always equals 30 exactly — every request either finished properly or was rejected cleanly, and nothing was abandoned mid-flight.

Common mistakes

No maximum wait on the drain. A single stuck request — waiting on a hung downstream call — can block a shutdown forever without one. Pair draining with a hard timeout, after which the process exits regardless.

The load balancer does not know about the drain. If traffic keeps arriving from outside after this server marks itself as draining, every one of those requests gets rejected pointlessly. The load balancer needs to stop routing here before the drain begins, usually via a health check this server starts failing on purpose.

Confusing SIGKILL with SIGTERM. On Linux, SIGTERM is the polite "please shut down" signal a drain step listens for. SIGKILL cannot be caught or handled at all — it terminates immediately. A deployment tool that gives too short a grace period before escalating to SIGKILL defeats the whole point of draining.

Testing shutdown only in a clean, idle state. The behaviour that matters is what happens under load, mid-request. Test by sending traffic and shutting down while it is still arriving, not by shutting down a server that was already quiet.

Try it yourself

Change DRAIN_AT to 0.1 — a shutdown signal that arrives almost immediately. Watch rejected climb close to 30, because almost no requests have had a chance to start before the drain begins.

What to learn next

Researcher — Mathematics and papers.

Kubernetes' shutdown sequence, precisely

When a Pod is deleted, Kubernetes does not send SIGTERM and immediately remove the Pod from service. The sequence is: the Pod is marked Terminating and removed from Endpoints (so Services stop routing to it) in parallel with SIGTERM being sent to the container. A terminationGracePeriodSeconds window (30 seconds by default) follows, during which the process is expected to drain. If the process has not exited by the end of that window, Kubernetes sends SIGKILL.

The subtlety that catches people: removal from Endpoints and delivery of SIGTERM are not strictly ordered relative to each other, because they happen through different subsystems (kube-proxy/endpoint controllers versus the kubelet). A container that stops accepting new work the instant it receives SIGTERM, before confirming it has been removed from Endpoints, can reject requests that were already in flight from the load balancer's perspective. The standard mitigation is a short unconditional sleep — a preStop hook — before the process begins refusing new work, giving the Endpoints removal time to propagate.

Connection draining at the load balancer layer

Cloud load balancers implement the load-balancer half of this independently — AWS ELB/ALB calls it "connection draining" or "deregistration delay"; GCP calls it "connection draining timeout." Once a backend is marked unhealthy or deregistering, the load balancer stops sending it new connections, but keeps existing connections open until they complete or a configured timeout elapses. This is the same two-phase pattern as the application-level drain above, applied one layer up the stack. Both layers are needed: an application that drains perfectly still fails users if the load balancer keeps sending it fresh traffic throughout.

Formalising the trade-off

Let $g$ be the grace period and $T$ the distribution of individual request durations. The probability a shutdown completes cleanly, without forcibly killing an in-flight request, is $P(\max(\text{in-flight durations}) \le g)$ — which shrinks as either the number of concurrent in-flight requests grows or the tail of $T$ gets heavier. This is precisely why services with occasional very slow requests — long-running LLM generations, large batch jobs — need a materially longer grace period than services with uniformly fast ones. A fixed default grace period is a guess that fits some workloads and badly under-serves others.

What to learn next