Scaling and Traffic Management

Serving from more than one region

Running your service from more than one place cuts the distance data has to travel, and keeps you serving even if one place goes down.

On this page 8
  1. The short answer
  2. The analogy you have already lived
  3. Why it exists
  4. How it works
  5. A real example you have seen
  6. The honest part
  7. Remember this
  8. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

The short answer

Multi-region serving means running your service from more than one physical location. Users then reach a nearby copy, instead of one far away.

The analogy you have already lived

You have noticed that a chai stall in your own neighbourhood serves you faster than one across the city ever could. The tea is not different. The distance is shorter. A chain with stalls in several neighbourhoods gives everyone a nearby option. Nobody has to travel across the whole city to one central stall.

If that one central stall runs out of gas for the evening, every customer in the city is stuck. A stall in each neighbourhood means the others keep serving, even while one is down.

Why it exists

Physical distance costs real time. A request from Mumbai to a server in Virginia has to travel there, and the answer has to travel back. That round trip is called latency. It is bounded by the speed of light, not by how fast your code runs. No amount of clever programming makes a signal cross the ocean faster.

A region is a distinct geographic location, typically an entire data centre or cluster of them. Running a copy of your service in more than one region puts a nearby copy within reach of more users. It also gives you somewhere to fail over to, if one region has a genuine outage.

How it works

   India user     -> nearest healthy region -> Mumbai   (short trip)
   Europe user    -> nearest healthy region -> Frankfurt (short trip)
   US-East user   -> nearest healthy region -> Virginia  (short trip)

   Mumbai goes down:

   India user     -> next nearest healthy region -> Frankfurt (longer trip)
   Europe user    -> unaffected -> Frankfurt
   US-East user   -> unaffected -> Virginia

Two separate benefits come out of the same setup. Everyday latency improves, from routing each user to somewhere physically near them. Failover improves too, from having somewhere else to route to when one region genuinely fails.

A real example you have seen

A video-streaming app starting instantly, regardless of which country you are in, relies on this. Nobody watching from Mumbai is quietly being served from a data centre in another continent for every request. That would be noticeably slower, and it is not what happens.

The honest part

Multi-region serving is expensive and genuinely complex. This lesson can only teach the reasoning behind it, using realistic example numbers, not real network measurement. Actually observing the effect needs real infrastructure in real geographic locations. That is beyond what any single machine can show you. Treat the numbers below as illustrative, not as a benchmark to trust blindly.

Remember this

  • Distance costs real time — physics, not code quality, sets the floor on cross-region latency.
  • Multi-region serving buys two things at once: everyday speed, from proximity, and failover, from redundancy.
  • It is genuinely expensive to run and operate. It becomes worth it once enough users, in enough different places, depend on the service.

What to learn next

Developer — Code and libraries.

Setup

No installs needed beyond the standard library.

Modelling routing and failover with realistic distances

These latency figures are illustrative, typical of real-world geography, not a live measurement. No code here can measure a real network without one actually existing between real data centres.

multi_region.py
# Illustrative round-trip latencies in milliseconds, typical of real geography.
# Not a live measurement: real numbers need a real network to observe.
LATENCY_MS = {
    "India":   {"Mumbai": 15,  "Frankfurt": 120, "Virginia": 220},
    "Europe":  {"Mumbai": 115, "Frankfurt": 12,  "Virginia": 90},
    "US-East": {"Mumbai": 230, "Frankfurt": 95,  "Virginia": 18},
}
USERS_PER_REGION = {"India": 500, "Europe": 300, "US-East": 300}


def best_region(user_region, down=frozenset()):
    options = {r: ms for r, ms in LATENCY_MS[user_region].items() if r not in down}
    return min(options, key=options.get)


def average_latency(down=frozenset()):
    total_weighted_ms = 0
    total_users = 0
    for user_region, count in USERS_PER_REGION.items():
        server_region = best_region(user_region, down)
        ms = LATENCY_MS[user_region][server_region]
        total_weighted_ms += ms * count
        total_users += count
        print(f"  {user_region:8s} -> {server_region:9s} ({ms} ms)")
    return total_weighted_ms / total_users


print("normal operation, every region healthy:")
normal_avg = average_latency()
print(f"  weighted average latency: {normal_avg:.1f} ms\n")

print("Mumbai region goes down, traffic fails over:")
failover_avg = average_latency(down={"Mumbai"})
print(f"  weighted average latency: {failover_avg:.1f} ms")
india_after = LATENCY_MS["India"][best_region("India", {"Mumbai"})]
print(f"  increase: {failover_avg - normal_avg:+.1f} ms overall,"
      f" but India users alone went from 15 ms to {india_after} ms")
Output
normal operation, every region healthy:
  India    -> Mumbai    (15 ms)
  Europe   -> Frankfurt (12 ms)
  US-East  -> Virginia  (18 ms)
  weighted average latency: 15.0 ms

Mumbai region goes down, traffic fails over:
  India    -> Frankfurt (120 ms)
  Europe   -> Frankfurt (12 ms)
  US-East  -> Virginia  (18 ms)
  weighted average latency: 62.7 ms
  increase: +47.7 ms overall, but India users alone went from 15 ms to 120 ms

Purely arithmetic over fixed, hand-written numbers — this reproduces exactly on any machine, but the input numbers themselves are illustrative, not measured.

Walking through it

best_region picks whichever server region has the lowest latency for a given user region, excluding any region marked down. This is a simplified stand-in for real geo-DNS or anycast routing, which makes the equivalent decision using real measured network paths rather than a hand-written table.

India users alone see the largest impact when Mumbai fails — from 15 ms to 120 ms, an 8x jump. Europe and US-East users notice nothing at all, because their nearest healthy region never changed. A regional outage does not hurt every user equally; it concentrates almost entirely on whoever was closest to the region that failed.

The overall weighted average — 62.7 ms — understates the real experience for India's users. It blends a large, badly affected group with two entirely unaffected groups. This is the same trap flagged in the researcher section of model serving: an average across very different experiences describes nobody's experience precisely.

Common mistakes

Assuming multi-region serving alone gives you failover, without also replicating data. Routing traffic to Frankfurt during a Mumbai outage only helps if Frankfurt actually has the data and the model that Mumbai had. Region failover is a routing problem plus a data-replication problem. Solving only the first one leaves you routing users to a region that cannot actually answer their request correctly.

Ignoring the cost of keeping regions in sync. Every region running a model needs that model's latest version. A new model deployed to Mumbai but forgotten in Frankfurt means users get different answers depending on region alone — a correctness problem hiding inside an infrastructure decision.

Routing purely on geographic distance, ignoring current load. The nearest region is not automatically the best choice if it happens to be overloaded right now. Real systems combine proximity with the load-aware routing from load balancing inference traffic, not proximity alone.

Try it yourself

Add a fourth region, "Singapore", with plausible latencies to each user region, and rerun average_latency(down={"Mumbai"}). Compare India's failover latency with Singapore available against Frankfurt alone, to see how much a well-placed extra region can soften a single region's outage.

What to learn next

Researcher — Mathematics and papers.

What actually routes users to a region

DNS-based geo-routing resolves a domain name to a different IP address, depending on the resolver's inferred location. It is simple to deploy, but caches at every layer between the user and the DNS server. Failover is only as fast as the lowest TTL respected along that chain — often minutes, not seconds.

Anycast advertises the same IP address from every region simultaneously, and relies on internet routing (BGP) to send each packet to the topologically nearest region automatically. Failover is near-instant, because it happens at the routing layer, not by waiting for a DNS record to expire — the mechanism behind services like Cloudflare's and Google's global load balancers.

Both approaches route on network proximity, which correlates with but is not identical to the geographic distance used for illustration above — real routing paths bend around undersea cable routes, peering agreements and congestion, not straight-line distance.

Consistency across regions

Once a model or its data is replicated across regions, a question arises immediately: what happens when two regions disagree, even briefly? Strong consistency requires every region to agree before any read succeeds — safe, but adds cross-region latency to every operation, defeating much of the purpose of having nearby regions. Eventual consistency allows brief disagreement, resolving it asynchronously — fast, but requires the application to tolerate a stale read occasionally. Most multi-region ML serving systems choose eventual consistency for model artefacts (a slightly stale model version for a few minutes is rarely catastrophic) and something closer to strong consistency for anything transactional, like billing.

Failover strategies, compared

  • Active-active: every region serves live traffic simultaneously; failover means routing around the failed one. Highest cost, lowest failover time.
  • Active-passive: one region serves traffic; others stand by, replicated but idle, ready to take over. Lower steady-state cost, slower failover — the passive region typically needs to be promoted and warmed.
  • Pilot light: a minimal, low-cost standby that can be scaled up during a real failure. It trades a longer failover time for a much lower idle cost, closely related to the scale to zero trade-off, applied at the region level rather than the instance level.

Reference

  • Google Cloud, Global load balancing overview — a detailed public account of anycast-based geo-routing at large scale, including how health checks drive failover between regions.

What to learn next