Contents

Architecture & System Design › System Design Fundamentals

Multi-Region Architecture

Running in several regions for lower latency and resilience.

Also known as: multi-region architecture, active-active regions, global architecture

Multi-region architecture serves users from several geographic regions — for latency (nearby serving), resilience (surviving regional failure), and compliance (data residency). Designs span active-passive (one serves, others stand by) to active-active (all serve, data replicates or partitions), with DNS/anycast steering and data strategy as the hard parts.

users → nearest region (serve) ; data → replicate / partition by residency
region fails → traffic shifts (designed, tested failover)

The cost ladder climbs steeply: multi-region reads (easy, replicas/CDN), multi-region writes (conflicts, coordination), and compliance-bounded data (residency partitions that can’t merge). Most systems should climb only as far as latency, resilience and law actually demand.

The classic mistakes:

  • Active-active without conflict design. Multi-region writes need resolution (CRDTs, LWW with eyes open, partitioned ownership) — discovered at the first split-brain otherwise.
  • Untested failover. DNS shifts, capacity elsewhere, data lag — regional failover fails in exactly the ways never rehearsed. Game-day the region loss, not just the node loss.
  • Data gravity ignored. Moving compute multi-region while data stays single-region buys latency nowhere and complexity everywhere. Data strategy leads; compute follows.
  • Residency as an afterthought. GDPR-style boundaries drawn through a globally-replicated store require painful surgery. Partition residency from the start where law demands it.
  • Capacity elsewhere assumed. Failover regions need headroom for the surge — idle capacity or fast autoscaling with tested limits. Hope is not headroom.
  • Consistent-everywhere expectations. Cross-region strong consistency costs hundreds of milliseconds per write; scope it to the operations truly needing it.
  • Operational multiplication. Every region multiplies deploys, monitors, incidents and on-call context. Count the human cost alongside the latency win.

How to adopt: read-globally first (CDN, replicas), write-locally where possible, go active-active only with conflict and residency designs proven. Regions are easy to draw and expensive to run — expand deliberately.