ExtraSysWork out the decision first. The vocabulary follows.
A finger touches an update button on a screen listing available, not reserved rows

FORK 04/When it breaks?/What an Availability Zone Actually Is

What an Availability Zone Actually Is

By the ExtraSys desk · When it breaks? · 4 min read

The term is everywhere in cloud architecture diagrams. What it means in practice — and what it quietly leaves unprotected — is less often explained.

The Physical Reality Behind the Label

An Availability Zone is a distinct, operationally isolated location within a cloud region. Each major hyperscaler implements the concept differently in its specifics, but the common intent is consistent: one AZ should be able to fail without taking the others down. AWS, Google Cloud, and Azure all publish this framing; the engineering underneath it is what matters.

In practice, an AZ maps to one or more discrete datacentre buildings. Those buildings share a region — meaning they are geographically close enough to make low-latency synchronous replication feasible, typically within a metropolitan area — but they are fed by separate power grids, separate cooling infrastructure, and separate physical network ingress. The separation is real, not virtual. A backhoe cutting a conduit, a transformer failure, or a flooding event that affects one AZ's facility should leave the others unaffected. Should. The engineering is designed for that outcome, not guaranteed to produce it.

The network connectivity between AZs inside a region uses dedicated, redundant paths that the provider owns and operates. Latency is low enough — single-digit milliseconds in most cases — that a database primary in one AZ and a synchronous replica in another is a practical architecture, not a theoretical one. This is the core design proposition: redundancy without the latency tax of spanning continents.

What AZs Do Not Protect Against

The isolation boundary is physical, not logical. This is where many architecture diagrams go quietly wrong.

A server rack aisle with dense cable management in blues and yellows, shot from floor level looking toward a hot-aisle containment door

The control plane — the APIs through which you provision instances, modify security groups, update load balancers, create snapshots — is typically a regional service, not a per-AZ one. When that control plane degrades, operations across all AZs in the region can fail simultaneously even when the underlying compute and storage in each AZ are still running. You may have healthy instances that you can no longer reach through the management layer, auto-scaling policies that stop responding, and load balancer configurations that cannot be updated. The data plane keeps running; the control plane does not. This distinction matters enormously for automated recovery logic that relies on API calls to reroute traffic or spin up replacements.

Shared regional services create the same exposure. A managed database service, a global load balancer endpoint, an identity and access management API — these are often regional constructs sitting above the AZ layer. Incidents that have affected major providers in the past have followed exactly this pattern: the zonal infrastructure was intact while a regional service degraded in ways that caused widespread customer impact. The shared-fate problem — the way logically distinct resources can share undisclosed dependencies — is the structural reason why AZ distribution, by itself, is not the same as resilience.

There is also the question of what happens inside a zone boundary when a subtle failure mode rather than a total outage occurs. A storage performance degradation, a network path with elevated packet loss, a hypervisor issue affecting a subset of hosts — these are partial failures, not clean zone outages. They can be harder to detect and harder to route around than a complete AZ loss. Architectural resilience that only accounts for the clean-cut scenario is incomplete.

What Good AZ Architecture Actually Looks Like

Distributing workloads across AZs is necessary but not complete. A few harder truths follow from the constraints above.

This distinction matters enormously for automated recovery logic that relies on API calls to reroute traffic or spin up replacements.

Stateless tiers are straightforward to distribute. Run compute across three AZs behind a load balancer, and an AZ loss typically causes a traffic redistribution that your health checks and auto-scaling handle automatically — assuming the control plane is functioning. The friction comes with stateful workloads. Databases, caches, and message queues require explicit decisions about replication, leader election, and what consistency you are prepared to sacrifice under partition. "Multi-AZ" on a managed service means the provider handles that replication automatically; it does not mean the service is immune to regional control-plane events.

Dependency mapping is underappreciated work. For each service your application calls, knowing whether it is zonal or regional — and what degrades when the regional layer is unavailable — is the difference between an architecture that survives zonal failure and one that merely looks like it does on a diagram. This is less glamorous than drawing arrows between AZ boxes, but more useful.

The appropriate mental model: AZs give you meaningful protection against physical infrastructure failures in a single location. They do not give you protection against the shared regional services that sit above them, the control plane that manages them, or failure modes that are partial rather than total. Those gaps require additional layers — or, for some workloads, a genuinely multi-region deployment whose own costs and complexity are worth understanding before committing to it.