Mainframe Path Start learning free
Applied11 min readLesson 3 of 3

Recovering consistently and testing DR

A DR site with the right disks is not yet a working business. Recovery needs a consistent restart of subsystems and batch, clear roles and runbooks, and regular realistic tests that expose the gaps before a real disaster does.

Restarting from a crash-consistent copy

With good replication, the recovery site has disks that look as if the primary lost power at a single instant. z/OS is started (IPLed) from those disks, then the subsystems restart. DB2, IMS, CICS and MQ each use their logs to back out units of work that had not committed and to complete ones that had. This is the same restart they perform after any unplanned outage, which is one reason the approach works.

Typical recovery order (varies by site)
z/OS and sysplexIPL, couple datasets
Security and catalogsRACF, master and user catalogs
DatabasesDB2, IMS restart
MiddlewareMQ, CICS
Scheduler and batchreconcile state
Validatebusiness checks

The batch problem

Online transactions are usually the easy part because each one is a unit of work that either committed or did not. Batch is harder. A disaster at 02:30 may hit with dozens of jobs part-way through, some updating VSAM files or sequential datasets that have no log to roll back. After recovery you must decide, job by job, what state the data is in.

Roles and runbooks

A disaster is the worst time to work out who decides what. DR plans define roles in advance, with deputies:

RoleResponsibility
DR or crisis leadDeclares the disaster, makes the invoke decision, owns communication with executives and regulators
Technical recovery leadCoordinates infrastructure, storage and system programmers through the recovery steps
Subsystem ownersRestart and check DB2, IMS, CICS, MQ and networking
Operations and batchReconcile the scheduler, decide reruns and restarts, release batch
Application ownersValidate business data and confirm the service is usable

Each role works from a DR runbook that lists steps, checks and contacts. Runbooks must be available when the primary site is gone, so they are kept somewhere that does not depend on it, often printed copies as well as an offsite system.

Testing DR

An untested DR plan is a hope. Tests range from cheap to realistic:

Test typeWhat it provesDisruption
WalkthroughPeople know the plan and their roleNone
Isolated recovery testSystems can be started from a copy of DR data in an isolated environmentLow
Full failover testProduction can actually run at the recovery siteHigh; usually planned for weekends
Run from DR site for a periodThe site can sustain real business loadHigh, but the strongest evidence

Many sites test at least annually, and regulators in financial services often expect evidence of regular, realistic tests. A useful test measures actual recovery time and data loss against the RTO and RPO, and every gap becomes a tracked action, just as in a post-incident review.

Common DR failures

TRY IT YOURSELF

Which objective (give its three-letter abbreviation) does a DR test check when it measures how long it took for customers to be able to pay again?

Show a hint

One objective is about data, the other about time to recover service.

Show the solution

RTO, the recovery time objective, is the maximum acceptable time to restore the service.

Common mistakes

Releasing batch straight after recovery

The scheduler's view may not match the data. Reconcile which jobs actually completed before releasing anything.

Keeping the DR plan only at the primary site

If the site is gone, so is the plan. Store runbooks and contact lists where they survive the disaster.

Testing only the easy part

Tests that stop once z/OS is IPLed prove little. Include subsystems, batch reconciliation, networks and business validation, and measure against RTO and RPO.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Replication, backups, tape and GDPSBack to Disaster recovery and business continuity