ExtraSysWork out the decision first. The vocabulary follows.
The Shared-Fate Problem

FORK 04/When it breaks?/The Shared-Fate Problem

The Shared-Fate Problem

By the ExtraSys desk · When it breaks? · 4 min read

Distributing your workloads across zones and regions feels like redundancy. It often isn't.

The Illusion of Independence

The pitch is tidy: deploy across multiple availability zones, and if one fails, the others absorb traffic. Deploy across multiple regions, and you survive a regional catastrophe. Both statements are technically defensible. Neither is the whole story.

The gap lives in what "independence" actually means at the infrastructure level. Availability zones within a region share physical proximity — typically within a metropolitan area — and more importantly, they share control-plane infrastructure. The control plane is the layer that manages, configures, and orchestrates your resources: the APIs you call to launch instances, scale groups, modify load balancers, update routing tables. It is, almost universally, regional, not zonal. When it degrades, every zone in the region feels it simultaneously. Your instances may still be running. Your ability to change anything about them may be gone.

This is the shared-fate problem: blast radii that the architecture diagrams don't draw.

Where Independence Breaks Down

Consider what actually happened during several well-documented hyperscaler outages in the 2010s and early 2020s. A recurring pattern: the data plane — the workloads themselves — kept running in all zones, while the control plane became unavailable or inconsistent. Autoscaling stopped working. Load balancers couldn't be reconfigured. DNS updates stalled. Operators watching healthy-looking dashboards found themselves unable to respond to any changing condition. Redundancy without control is a frozen posture: you survive the first hit, but you can't adapt to the second.

The dependency graph goes deeper than control planes. Managed identity services, secret stores, certificate authorities, and service-mesh control planes are often implemented as regional singletons or near-singletons. An application that authenticates every request against a managed identity provider inherits that provider's availability profile, regardless of how many zones its compute spans. The same logic applies to managed databases with a regional primary, centralised log aggregators, and configuration services. Each of these is a potential choke point that cuts across the zone topology the rest of the architecture is supposed to make redundant.

DNS is its own chapter. A misconfiguration or propagation failure in a regional DNS resolver can make zone-distributed services unreachable even when the services themselves are fully operational. This isn't hypothetical — it's a recurring cause in cloud postmortem literature, and it's a reminder that the network path to your redundant infrastructure carries its own failure modes.

Cross-zone traffic also introduces a subtler problem: the assumption that zone-to-zone bandwidth is effectively unlimited. It isn't, and under stress — when one zone has failed and the surviving zones absorb the full load — internal bandwidth can saturate in ways that normal load patterns never reveal. Stress tests that simulate a full-zone failure are uncommon. When they are run, the results are frequently surprising.

A close-up of a fibre optic patch panel with labelled LC connectors in a grey steel chassis

The Multi-Region Mirage

Lifting the scope to multi-region doesn't eliminate shared fate; it displaces it. Multi-region is harder than it sounds precisely because a genuine multi-region architecture requires solving data-consistency, conflict-resolution, and failover-coordination problems that most teams underestimate until they're standing in the middle of an incident. But beyond those application-layer problems, there are infrastructure-layer shared dependencies that persist even in a well-designed multi-region setup.

Global load balancing services, edge networks, and provider-managed CDN tiers are often operated at a scope that crosses regional boundaries. A bug in the provider's global routing infrastructure — not a regional failure, but a software defect in a globally-deployed control system — can affect multiple regions simultaneously. This is not speculation: it is the category of incident responsible for some of the most widely noticed cloud outages in recent years. The provider's postmortem will describe it as a "global" or "multi-region" event and will note that no single AZ or region was the cause.

Shared BGP peering arrangements, provider-backbone congestion, and anycast misbehaviour follow the same pattern. They are failure modes that exist entirely above the region/zone abstraction and are therefore invisible to architectures designed around that abstraction.

Consider what actually happened during several well-documented hyperscaler outages in the 2010s and early 2020s.

What This Means Practically

None of this is an argument against zone distribution. It remains worth doing: it protects against the most common class of failure, which is the isolated hardware or facility event. But it should be understood for what it is — a partial mitigation with a defined blast radius, not a guarantee.

The honest design exercise is to map the actual dependency graph, not the compute topology. For each component — identity, secrets, DNS, database, configuration, service mesh — ask where the singleton lives and what happens to your application when it becomes unavailable or inconsistent. Then ask whether your runbook assumes control-plane availability, because if it does, you may be planning your recovery using tools that don't work during the event you're recovering from.

Understanding what an availability zone actually is — physically and logically — is the prerequisite. The AZ model is a real and useful guarantee about certain kinds of isolation. It is not a guarantee about the shared systems that sit above and beside it. Knowing where that guarantee stops is the beginning of an honest redundancy design.