Mainframe Path Start learning free
Applied10 min readLesson 1 of 3

RPO, RTO and business impact

Disaster recovery starts with two business questions: how much data can we afford to lose, and how long can we be down? A business impact analysis answers them per service, and the answers drive every technical choice and its cost.

Continuity, recovery and availability

Business continuity is the organisation's ability to keep delivering critical services during and after a disruption, including people, buildings, suppliers and communications. Disaster recovery (DR) is the IT part of that: getting systems and data running again, usually at another site, after something takes out the primary. High availability is related but different: it keeps a service running through smaller failures, such as one LPAR or one CICS region going down, often using a Parallel Sysplex in one site.

Three layers of protection
Business continuitypeople, premises, suppliers, communications
Disaster recoverysystems and data at an alternate site
High availabilitysurvive component failures without an outage

The two numbers that matter

TermQuestion it answersExample
RPO: recovery point objectiveHow much data, measured in time, can we lose?RPO of 5 minutes: after recovery, at most the last 5 minutes of updates may be missing
RTO: recovery time objectiveHow long until the service must be running again?RTO of 2 hours: customers must be able to pay within 2 hours of the decision to recover
RPO looks backwards from the disaster, RTO looks forwards
Last good copydata is safe up to here
RPO gapupdates at risk
Disasterprimary lost
RTO gapservice unavailable
Service restoredat recovery site

An RPO of zero means no committed data may be lost, which in practice needs synchronous replication. An RTO measured in minutes needs a recovery site that is already running and highly automated. Both cost far more than an RPO of 24 hours restored from last night's backups, so they should only be bought where the business genuinely needs them.

Business impact analysis

A business impact analysis (BIA) is how the business decides those numbers. For each business process it asks what happens as an outage grows: financial loss, regulatory breach, customer harm, reputational damage. It also maps which IT services the process depends on, and what they in turn depend on.

  1. List business processes and their owners.
  2. Estimate impact over time: after 1 hour, 4 hours, 1 day, 1 week.
  3. Set RTO and RPO per process based on where impact becomes unacceptable.
  4. Map dependencies down to applications, databases, batch, networks and third parties.
  5. Rank recovery order and check it against what the infrastructure can deliver.

Dependencies are where mainframe DR gets subtle. A mobile payment might pass through an API gateway, z/OS Connect, CICS, DB2 and MQ, plus a fraud service on another platform. Recovering the mainframe in 30 minutes is pointless if the network routes or the distributed fraud service take six hours.

DR tiers at concept level

A long-standing way to describe DR capability is the set of tiers first defined by the SHARE user group and widely used in IBM material. The exact wording varies between sources, but the progression is consistent:

TierRough descriptionTypical RPO / RTO
0No offsite dataPossibly unrecoverable
1-2Backups shipped offsite; tier 2 adds a standby (hot) siteUp to a day or more / days
3Electronic vaulting: backups sent over the networkHours / a day or more
4Point-in-time copies, active secondary siteHours / hours
5Transaction integrity across both sitesSeconds to minutes / hours
6Continuous replication, little or no data lossNear zero / minutes to hours
7Replication plus automated, site-wide failoverNear zero / minutes to an hour

The ranges are indicative only; your site's documented capability and test results are what count. Tiers are useful to describe roughly where a service sits and what moving up would involve.

What counts as a disaster

Traditional planning focused on losing a building: fire, flood, power or network failure. Today the scenario many regulators emphasise is cyber: ransomware or a malicious insider corrupting data. That changes the picture, because a corruption is replicated to the DR site just as faithfully as a good update. Lesson 2 returns to this.

Common mistakes

Treating RPO and RTO as the same thing

RPO is about data loss looking back; RTO is about downtime looking forward. A design can meet one and badly miss the other.

Setting objectives without the business

IT can describe options and costs, but only the business can say what loss and downtime are acceptable. Base them on a BIA with named owners.

Ignoring non-mainframe dependencies

A recovered mainframe does not restore a service whose network, gateway or distributed components are still down. Map end-to-end dependencies.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
Replication, backups, tape and GDPS →