Fork 04
When it breaks?
Availability zones, regions, failure modes and what the architecture actually protects against — and what it does not.
work down the list until one of them stings
The worksheet
Every guide at this fork answers one row. Work down the left column until something you believe stops being true.
| What you think | What is actually true |
|---|---|
| Multi-zone means safe. | A zone buys distance from a building. It buys nothing at all against your own deploy.What an Availability Zone Actually Is → |
| Spread across zones means independent. | Shared control planes and common dependencies cross zone boundaries without asking.The Shared-Fate Problem → |
| We have an RTO; it is in the runbook. | A target that has never been tested is an ambition. The gap between the two is where data actually dies.RTO and RPO Without the Jargon → |
| Multi-region is the obvious next step up. | Consistency, latency and operational complexity have costs that can exceed the risk being mitigated.Why Multi-Region Is Harder Than It Sounds → |
| The postmortem explains what happened. | It discloses some things and consistently omits others. Reading one properly is a skill.The Anatomy of a Cloud Outage Postmortem → |
The 5 guides, in reading order
Every guide →
What an Availability Zone Actually IsAn explainer of the physical and network isolation that AZs are designed to provide — and the failure modes they explicitly do not protect against. Covers the distinction between zone failure and regional failure without invented SLA figures.By the ExtraSys desk · When it breaks? · 4 min read
The Shared-Fate ProblemAn argument about why distributing workloads across AZs or regions does not guarantee independence — shared control planes, common dependencies, and blast-radius patterns that cross zone boundaries. Sceptical, evidence-reasoned.By the ExtraSys desk · When it breaks? · 4 min read
RTO and RPO Without the JargonA plain-language explainer of recovery time and recovery point objectives — what they measure, how they are set, and why stated targets are often not tested. Short and useful for decision-makers who have seen the terms but not been given honest definitions.By the ExtraSys desk · When it breaks? · 2 min read
Why Multi-Region Is Harder Than It SoundsA long argument that multi-region deployment is often proposed as a solution without accounting for the data-consistency, latency, and operational complexity costs it introduces — and that those costs can exceed the risk being mitigated.By the ExtraSys desk · When it breaks? · 6 min read
The Anatomy of a Cloud Outage PostmortemA structural walk-through of how major cloud providers publish postmortems — what they disclose, what they consistently omit, and how to read between the lines to understand actual failure causation.By the ExtraSys desk · When it breaks? · 4 min read