
FORK 04/When it breaks?/RTO and RPO Without the Jargon
RTO and RPO Without the Jargon
Two numbers sit at the heart of every recovery plan. Most organisations have chosen them without testing whether they are achievable.
What the numbers actually mean
RTO — Recovery Time Objective — is the maximum acceptable duration between a failure and restored service. If your RTO is four hours, you are saying that four hours of downtime is tolerable. RPO — Recovery Point Objective — is the maximum acceptable data loss, expressed as time: an RPO of one hour means you can afford to lose up to an hour's worth of transactions.
Neither number is a capability. Both are statements of business tolerance, set by asking what downtime actually costs — lost revenue, regulatory exposure, contractual penalties, reputational damage — and what data loss costs in concrete terms. That conversation is harder than it sounds, and many organisations skip it, inheriting round numbers from a template or a vendor's defaults.

The engineering challenge is then to build a recovery architecture whose real performance meets those targets. Replication frequency determines RPO; provisioning speed and failover automation determine RTO. Smaller numbers in both directions cost more: continuous replication is more expensive than nightly backups, and a warm standby that can take traffic in minutes costs more than a cold archive that takes hours to restore.
The most common failure mode is not a gap in the plan — it is a plan that has never been tested under realistic conditions. Backup jobs that run cleanly do not guarantee restores that succeed at speed. A stated two-hour RTO that has never been rehearsed is, in practice, aspirational fiction.
Set the numbers from business requirements, not from what your current infrastructure happens to do. Then test the recovery — not the backup — at least annually, under conditions that resemble an actual incident. The gap between stated and demonstrated is where real risk lives.