RPO, RTO and business impact
Disaster recovery starts with two business questions: how much data can we afford to lose, and how long can we be down? A business impact analysis answers them per service, and the answers drive every technical choice and its cost.
Continuity, recovery and availability
Business continuity is the organisation's ability to keep delivering critical services during and after a disruption, including people, buildings, suppliers and communications. Disaster recovery (DR) is the IT part of that: getting systems and data running again, usually at another site, after something takes out the primary. High availability is related but different: it keeps a service running through smaller failures, such as one LPAR or one CICS region going down, often using a Parallel Sysplex in one site.
The two numbers that matter
| Term | Question it answers | Example |
|---|---|---|
| RPO: recovery point objective | How much data, measured in time, can we lose? | RPO of 5 minutes: after recovery, at most the last 5 minutes of updates may be missing |
| RTO: recovery time objective | How long until the service must be running again? | RTO of 2 hours: customers must be able to pay within 2 hours of the decision to recover |
An RPO of zero means no committed data may be lost, which in practice needs synchronous replication. An RTO measured in minutes needs a recovery site that is already running and highly automated. Both cost far more than an RPO of 24 hours restored from last night's backups, so they should only be bought where the business genuinely needs them.
Business impact analysis
A business impact analysis (BIA) is how the business decides those numbers. For each business process it asks what happens as an outage grows: financial loss, regulatory breach, customer harm, reputational damage. It also maps which IT services the process depends on, and what they in turn depend on.
- List business processes and their owners.
- Estimate impact over time: after 1 hour, 4 hours, 1 day, 1 week.
- Set RTO and RPO per process based on where impact becomes unacceptable.
- Map dependencies down to applications, databases, batch, networks and third parties.
- Rank recovery order and check it against what the infrastructure can deliver.
Dependencies are where mainframe DR gets subtle. A mobile payment might pass through an API gateway, z/OS Connect, CICS, DB2 and MQ, plus a fraud service on another platform. Recovering the mainframe in 30 minutes is pointless if the network routes or the distributed fraud service take six hours.
DR tiers at concept level
A long-standing way to describe DR capability is the set of tiers first defined by the SHARE user group and widely used in IBM material. The exact wording varies between sources, but the progression is consistent:
| Tier | Rough description | Typical RPO / RTO |
|---|---|---|
| 0 | No offsite data | Possibly unrecoverable |
| 1-2 | Backups shipped offsite; tier 2 adds a standby (hot) site | Up to a day or more / days |
| 3 | Electronic vaulting: backups sent over the network | Hours / a day or more |
| 4 | Point-in-time copies, active secondary site | Hours / hours |
| 5 | Transaction integrity across both sites | Seconds to minutes / hours |
| 6 | Continuous replication, little or no data loss | Near zero / minutes to hours |
| 7 | Replication plus automated, site-wide failover | Near zero / minutes to an hour |
The ranges are indicative only; your site's documented capability and test results are what count. Tiers are useful to describe roughly where a service sits and what moving up would involve.
What counts as a disaster
Traditional planning focused on losing a building: fire, flood, power or network failure. Today the scenario many regulators emphasise is cyber: ransomware or a malicious insider corrupting data. That changes the picture, because a corruption is replicated to the DR site just as faithfully as a good update. Lesson 2 returns to this.
Common mistakes
RPO is about data loss looking back; RTO is about downtime looking forward. A design can meet one and badly miss the other.
IT can describe options and costs, but only the business can say what loss and downtime are acceptable. Base them on a BIA with named owners.
A recovered mainframe does not restore a service whose network, gateway or distributed components are still down. Map end-to-end dependencies.
What you will see at work
- Application teams are asked each year to confirm the RPO, RTO and dependencies of their services for the BIA.
- Financial regulators in many countries expect firms to define impact tolerances for important business services and prove they can stay within them.
- DR architecture discussions for the mainframe start from the most demanding RPO and RTO the platform must support.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.