ExtraSysWork out the decision first. The vocabulary follows.
The Anatomy of a Cloud Outage Postmortem

FORK 04/When it breaks?/The Anatomy of a Cloud Outage Postmortem

The Anatomy of a Cloud Outage Postmortem

By the ExtraSys desk · When it breaks? · 4 min read

What providers put in their incident reports, what they leave out, and how to read the gap.

What a Postmortem Actually Is

When a major cloud provider suffers a significant outage, a postmortem document eventually appears on a status page or engineering blog. The format varies by provider, but the structure is remarkably consistent: a timeline, a description of the triggering event, a root-cause section, and a list of corrective actions. These documents are real engineering products — not press releases written by lawyers, though lawyers have clearly read them. Understanding what the genre is designed to do helps you extract what it is designed to obscure.

The timeline is almost always accurate as a sequence of events. Timestamps are verifiable; they show up in monitoring alerts, customer support tickets, and third-party outage trackers. Where postmortems are most faithful is in what happened when. Where they are least faithful is in why it took so long to resolve, and who made which decision at which point.

The Standard Anatomy

A postmortem follows a predictable shape. Read enough of them and the sections become templates.

The triggering event is described with precision and usually involves something technically defensible — a configuration change, a capacity edge case, a dependency that behaved unexpectedly under load. This framing matters. It positions the failure as a systems problem rather than a judgment problem, which is both intellectually honest (complex systems do fail in non-obvious ways) and organisationally convenient (no one is identifiably responsible).

The root-cause section is where careful reading pays off. Providers typically use phrases like "an unexpected interaction between" or "a condition not anticipated by." These are not evasions exactly — distributed systems genuinely produce emergent failure modes — but they function to abstract away the human decisions upstream. A misconfiguration usually has an author. A capacity planning gap usually has a process owner. The postmortem rarely names either.

The impact section describes scope in the vaguest terms that remain defensible: affected regions, impaired services, percentage of requests failing. You will almost never see an estimate of customer-hours lost or a commercial impact figure. Providers have no obligation to publish these, and the liability exposure from specificity is obvious.

Corrective actions are where optimism blooms. The list is always longer than you'd expect, and the items are always reasonable — automated rollback improvements, better canary testing, enhanced monitoring thresholds. What you cannot verify from outside is which items are already closed, which are still open six months later, and whether any of the same failure modes reappear in a subsequent incident. Cross-referencing postmortems against later incidents for the same service is one of the more revealing exercises a reliability engineer can do.

A row of UPS units in a datacentre with indicator lights, shot from the side at mid-height

What Gets Left Out

Three categories of information are consistently absent.

First: the decision chain. When an incident drags on — when the initial mitigation makes things worse, or when the wrong rollback procedure is attempted — the postmortem describes the outcome without reconstructing who approved what and when. This is understandable from a legal standpoint and probably the right call for individual-blame reasons, but it makes it genuinely hard to assess whether the provider's incident-response culture is improving.

Second: dependency opacity. Cloud outages frequently cascade through internal dependencies that customers cannot see, instrument, or anticipate. When a provider's internal control plane is impaired, it can prevent customers from performing their own recovery actions — spinning up instances in a different zone, triggering a failover, modifying routing rules. Postmortems acknowledge this in passing but rarely expose the full dependency graph. As a reader, you should treat any outage that prevented customer-side remediation as a signal that shared responsibility breaks down in exactly the moment you most need it.

Third: recurrence patterns. Providers do not cross-reference their own published postmortems explicitly. If a class of failure — say, a configuration propagation bug — has appeared three times in five years, no individual postmortem will say so. That pattern-matching is work you have to do yourself, and it is some of the most useful reliability research available to any team evaluating a provider or designing a recovery architecture.

Check whether your own architecture would have shielded you from the blast radius described, or whether you were relying on the same control plane that failed.

How to Read One Productively

Approach postmortems as structured evidence, not verdicts. Strip the corrective-actions section and ask whether the underlying failure mode can recur before any of those actions complete — because in complex systems, it often can. Check whether your own architecture would have shielded you from the blast radius described, or whether you were relying on the same control plane that failed. Treat the timeline as the raw material for a tabletop exercise: could your team have detected and responded independently if the provider's status page had stayed green?

The best use of a cloud postmortem is not to assign blame to the provider. It is to find the places where your design assumed the provider would not fail, and then ask whether that assumption is load-bearing.