Latency, Load Testing and Capacity

Soak testing and slow memory leaks

A soak test runs a service under ordinary load for a long stretch of time, to catch slow problems — like memory quietly climbing — that a short test never runs long enough to see.

Read these first

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

A soak test runs a service for a long stretch of time, to catch problems that only show up slowly.

The analogy you have already lived

You have seen a dripping tap left overnight leave a full bucket by morning, when a five-second glance at the same tap showed nothing worth noticing. The drip was always there. It needed time to become visible.

Some software problems behave exactly like that drip — invisible in a quick check, obvious after hours.

Why it exists

Most testing checks whether something works right now. Send a request, get a correct answer, done. That kind of test genuinely proves correctness, and it proves nothing about what happens after the ten-thousandth request.

A memory leak is one of the classic problems that only a long test reveals. It is a program that keeps a little more memory reserved after every request, and never gives it back. Ten requests, or even a thousand, might use a completely unremarkable amount of memory. A hundred thousand requests, over hours, can quietly grow that until the server runs out of memory and crashes. It often happens overnight, or over a long weekend, far from anyone watching.

A soak test — sometimes called an endurance test — runs a service under steady, realistic load for an extended period specifically to catch this. It is not looking for a fast crash. It is looking for a slow one.

How it works

A short test (a few minutes):
   memory: |----|----|----|         looks flat, looks fine

The SAME leaking code, soak-tested for hours:
   memory: |----|----|----|----|----|----|----|----|~~~~ still climbing
                                                      eventually: crash

A healthy server, soak-tested the same way:
   memory: |----|----|----|----|----|----|----|----|----  stays flat

A real example you have seen

A phone or laptop that runs smoothly right after a restart, then gets sluggish after being left on for days, is often showing exactly this pattern. Something in the system is slowly holding onto memory it should have released.

The honest part

A memory leak rarely announces itself. The service keeps answering requests correctly the entire time it is leaking — right up until it runs out of memory. Correctness and health are not the same thing, and a soak test is one of the few tools that checks health specifically.

Remember this

  • A soak test runs a service under real load for a long time, to catch slow problems a short test cannot.
  • A memory leak is memory that keeps growing and never comes back down, even while every individual request still works.
  • The service can look completely healthy, right up until it suddenly is not.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install psutil

A real leak, and a real fix, measured over real memory

soak_test.py
import os
import psutil

process = psutil.Process(os.getpid())

REQUEST_LOG = []   # the bug: every request's payload is kept forever


def leaky_handler(request_id):
    # A common real mistake: logging full request/response objects into
    # an in-memory list "for debugging", and never trimming it.
    payload = {"id": request_id, "body": "x" * 2000}
    REQUEST_LOG.append(payload)
    return {"ok": True}


def fixed_handler(request_id, recent, max_keep=200):
    payload = {"id": request_id, "body": "x" * 2000}
    recent.append(payload)
    if len(recent) > max_keep:
        recent.pop(0)   # bounded: old entries are dropped
    return {"ok": True}


def rss_mb():
    return process.memory_info().rss / (1024 * 1024)


def soak(handler_name, fn, n_requests, checkpoint_every):
    print(f"-- {handler_name} --")
    readings = []
    for i in range(n_requests):
        fn(i)
        if (i + 1) % checkpoint_every == 0:
            mb = rss_mb()
            readings.append(mb)
            print(f"  after {i+1:6d} requests: RSS = {mb:7.1f} MB")
    return readings


if __name__ == "__main__":
    N = 60000
    STEP = 10000

    leaky_readings = soak("leaky handler (unbounded REQUEST_LOG)", leaky_handler, N, STEP)

    recent = []
    fixed_readings = soak("fixed handler (bounded to 200 entries)",
                           lambda i: fixed_handler(i, recent), N, STEP)

    print()
    print(f"leaky handler grew by {leaky_readings[-1]-leaky_readings[0]:6.1f} MB "
          f"over {N-STEP} requests after the first checkpoint")
    print(f"fixed handler grew by {fixed_readings[-1]-fixed_readings[0]:6.1f} MB "
          f"over {N-STEP} requests after the first checkpoint")
Output
-- leaky handler (unbounded REQUEST_LOG) --
  after  10000 requests: RSS =    20.4 MB
  after  20000 requests: RSS =    23.3 MB
  after  30000 requests: RSS =    25.9 MB
  after  40000 requests: RSS =    28.6 MB
  after  50000 requests: RSS =    31.1 MB
  after  60000 requests: RSS =    34.2 MB
-- fixed handler (bounded to 200 entries) --
  after  10000 requests: RSS =    34.3 MB
  after  20000 requests: RSS =    34.3 MB
  after  30000 requests: RSS =    34.3 MB
  after  40000 requests: RSS =    34.3 MB
  after  50000 requests: RSS =    34.3 MB
  after  60000 requests: RSS =    34.3 MB

leaky handler grew by   13.8 MB over 50000 requests after the first checkpoint
fixed handler grew by    0.0 MB over 50000 requests after the first checkpoint

This is a real, measured memory reading from this process's own operating-system memory usage — not a simulation. RSS, resident set size, is the actual physical memory this process holds, reported by the operating system itself through psutil. The leaky handler's memory climbs steadily, checkpoint after checkpoint. The fixed handler's memory, after its first checkpoint, does not move at all. Your own absolute numbers will differ by platform and Python version — the shape of a rising line versus a flat one is the real, trustworthy finding.

Line-by-line walkthrough

process.memory_info().rss asks the operating system directly how much physical memory this process is using right now — the same number a task manager or top would show.

leaky_handler appends every request to a module-level list that nothing ever removes from — this is the entire bug, and it is a completely realistic one: exactly the kind of debug-logging line that gets left in by accident.

fixed_handler does the same logging, but caps the list at max_keep entries with recent.pop(0), discarding the oldest entry once the cap is reached — bounded memory, regardless of how many requests arrive in total.

Common mistakes

Testing memory for a few seconds and declaring it fine. As the output above shows, a fixed handler is flat from its very first checkpoint. A leak needs enough volume and enough time to become visible against normal noise — a short check cannot tell the two apart.

Growing caches, logs, or connection pools with no upper bound. Any "keep everything, to be safe" data structure attached to a long-running process is a leak waiting to happen, whether or not it was intended as a cache.

Watching total system memory instead of the process's own RSS. Other processes on the same machine move total memory around too. Track the specific process you care about.

Assuming Python's garbage collector prevents this category of bug entirely. It does not — Python reliably frees memory that is no longer referenced, but a growing list that is still referenced, like REQUEST_LOG above, is never eligible for collection. The garbage collector is not a leak detector.

Try it yourself

Change max_keep from 200 to 2_000_000 — a cap so large it is effectively unbounded — and rerun the "fixed" handler. Its memory line should start climbing too, at a slower rate than the fully unbounded version, showing that a "bound" only protects you if it is actually small enough to matter.

What to learn next

Researcher — Mathematics and papers.

What RSS does and does not tell you

RSS (resident set size) counts physical memory pages currently mapped into the process, including shared library pages counted once per process even if shared across many. It differs from VSZ (virtual size, which can vastly overstate real usage) and from Python's own sys.getsizeof-style accounting, which measures object size in isolation and misses fragmentation, allocator overhead, and C-extension-owned memory entirely. For diagnosing "is this process's footprint growing," RSS sampled over time, as done above, is the most directly meaningful of the common options.

Distinguishing a leak from fragmentation

Not every memory growth pattern is a true leak — memory that is genuinely referenced and will eventually be freed. CPython's memory allocator (pymalloc) can also exhibit fragmentation: freed small objects leave gaps that are not returned to the operating system, inflating RSS without any object actually being unreachable-but-retained. tracemalloc (standard library) distinguishes these cases by tracking allocations back to the Python code that made them, via tracemalloc.take_snapshot() compared across two points in time — the standard technique for finding exactly which line is responsible for retained memory, as opposed to RSS's black-box "the number went up."

Reference-cycle leaks, specifically

A subtler leak class: objects that reference each other in a cycle (a.next = b; b.prev = a) are not freed by CPython's primary reference-counting mechanism, since neither object's count ever drops to zero on its own. The generational garbage collector (gc module) does eventually collect ordinary reference cycles — but an object defining __del__ in a cycle historically could defeat even that (fixed in CPython 3.4, PEP 442), and objects held by a C extension outside Python's own accounting can leak regardless of gc. gc.get_objects() and gc.collect()'s return count are the standard entry points for diagnosing this specific case, distinct from the unbounded-list pattern demonstrated above.

Soak test design in production

A production soak test typically runs realistic load for a duration meaningfully longer than any internal cache TTL, connection pool recycle interval, or scheduled maintenance window in the system under test — long enough that any periodic-but-slow accumulation has time to complete at least one full cycle and reveal whether it is bounded. Testing for a duration shorter than a leak's own periodicity can pass cleanly while masking a real problem, symmetrically to how a load test that is too short can miss tail latency effects that only appear over volume.

References

  • Python documentation, tracemalloc — docs.python.org/3/library/tracemalloc.html
  • PEP 442, Safe Object Finalization, 2013 — the fix allowing the cyclic garbage collector to reliably collect objects with __del__ methods.
  • Gregg, B., Systems Performance, 2nd ed., 2020, Chapter 7, Memory — a thorough treatment of RSS, virtual memory, and leak diagnosis at the operating-system level.

What to learn next