Recovering consistently and testing DR
A DR site with the right disks is not yet a working business. Recovery needs a consistent restart of subsystems and batch, clear roles and runbooks, and regular realistic tests that expose the gaps before a real disaster does.
Restarting from a crash-consistent copy
With good replication, the recovery site has disks that look as if the primary lost power at a single instant. z/OS is started (IPLed) from those disks, then the subsystems restart. DB2, IMS, CICS and MQ each use their logs to back out units of work that had not committed and to complete ones that had. This is the same restart they perform after any unplanned outage, which is one reason the approach works.
The batch problem
Online transactions are usually the easy part because each one is a unit of work that either committed or did not. Batch is harder. A disaster at 02:30 may hit with dozens of jobs part-way through, some updating VSAM files or sequential datasets that have no log to roll back. After recovery you must decide, job by job, what state the data is in.
- Which jobs had finished? The scheduler database must itself be recovered consistently with the data, or its view will not match reality.
- Which were running? Their outputs may be partly written. Check for checkpoint and restart records, and use the same rules as for any rerun or restart.
- Which inputs arrived? Files transmitted from partners just before the disaster may need to be resent.
- What about generation datasets? GDG relative numbers depend on the catalog; confirm the right generations are referenced.
Roles and runbooks
A disaster is the worst time to work out who decides what. DR plans define roles in advance, with deputies:
| Role | Responsibility |
|---|---|
| DR or crisis lead | Declares the disaster, makes the invoke decision, owns communication with executives and regulators |
| Technical recovery lead | Coordinates infrastructure, storage and system programmers through the recovery steps |
| Subsystem owners | Restart and check DB2, IMS, CICS, MQ and networking |
| Operations and batch | Reconcile the scheduler, decide reruns and restarts, release batch |
| Application owners | Validate business data and confirm the service is usable |
Each role works from a DR runbook that lists steps, checks and contacts. Runbooks must be available when the primary site is gone, so they are kept somewhere that does not depend on it, often printed copies as well as an offsite system.
Testing DR
An untested DR plan is a hope. Tests range from cheap to realistic:
| Test type | What it proves | Disruption |
|---|---|---|
| Walkthrough | People know the plan and their role | None |
| Isolated recovery test | Systems can be started from a copy of DR data in an isolated environment | Low |
| Full failover test | Production can actually run at the recovery site | High; usually planned for weekends |
| Run from DR site for a period | The site can sustain real business load | High, but the strongest evidence |
Many sites test at least annually, and regulators in financial services often expect evidence of regular, realistic tests. A useful test measures actual recovery time and data loss against the RTO and RPO, and every gap becomes a tracked action, just as in a post-incident review.
Common DR failures
- New volumes, applications or datasets added to production but not to replication or consistency groups.
- Expired passwords, certificates or licence keys that only matter at the recovery site.
- Hardware differences at the recovery site, such as fewer processors or different network addresses, that stop work starting or meeting WLM goals.
- Runbooks referring to people who have left or systems that no longer exist.
- Recovery of the mainframe tested alone while dependent distributed services were never included.
- No tested plan for coming back to the primary site afterwards.
Which objective (give its three-letter abbreviation) does a DR test check when it measures how long it took for customers to be able to pay again?
Show a hint
One objective is about data, the other about time to recover service.
Show the solution
RTO, the recovery time objective, is the maximum acceptable time to restore the service.
Common mistakes
The scheduler's view may not match the data. Reconcile which jobs actually completed before releasing anything.
If the site is gone, so is the plan. Store runbooks and contact lists where they survive the disaster.
Tests that stop once z/OS is IPLed prove little. Include subsystems, batch reconciliation, networks and business validation, and measure against RTO and RPO.
What you will see at work
- Operations and batch staff play a large role in DR tests, reconciling the schedule and restarting the batch cycle at the recovery site.
- DR tests often happen over weekends and involve dozens of teams following detailed runbooks and checklists.
- Findings from DR tests feed into problem and change management, for example adding new volumes to consistency groups.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.