Consistency, testing and cost trade-offs
A DR copy is only useful if batch, online and system data line up and the team has proven they can recover from it. This lesson covers recovering in-flight batch and online work, how to test DR without risking production, and how to weigh availability against cost.
Consistency beyond the disk
A consistency group gives a crash-consistent copy: as if every system had stopped at one instant. Subsystems such as DB2, IMS, MQ and CICS then do what they do after any crash: read their logs, back out uncommitted units of work and complete committed ones. Online transactions in flight are simply backed out, and users retry. Batch is harder.
In-flight batch
A batch job half-way through updating a sequential file or a GDG has no automatic backout. After failover, operations must know exactly which jobs were running, which had completed, and which datasets they had created. That information lives in the scheduler's database, in the catalogs and in the job's own checkpoint records, so all of them must be in the replicated, consistent set.
| Work type | State after failover | Recovery approach |
|---|---|---|
| Online DB2 / IMS / CICS | Uncommitted work backed out by subsystem restart | Automatic; users retry |
| Persistent MQ messages | Recovered from MQ logs | Automatic; check for duplicates if not under syncpoint |
| Batch updating databases with commits | Committed to last commit point | Restart from checkpoint |
| Batch writing sequential files or GDGs | Partial outputs may exist | Delete partial outputs, rerun step |
| File transfers to partners | Unknown whether delivered | Confirm with partner before resending |
Designing batch to be idempotent or restartable at step level makes this manageable. Jobs that update in place without checkpoints, or that delete their input before the output is safe, turn a DR event into a data-repair exercise.
Testing DR
An untested DR plan is a hope. Tests should prove three things: the data copy is usable, the people and runbooks work, and the RTO is achievable. Common test styles:
- Component test: bring up one subsystem from the replicated copy to check it starts cleanly.
- Isolated full test: take a point-in-time copy of the secondary disk (for example with FlashCopy or a vendor equivalent) and IPL systems from it on an isolated network, so replication to production continues during the test.
- Planned site switch: move production to the other site for real during a quiet period and run there. The strongest proof, with real risk.
- Tabletop exercise: walk the runbook with the team to find gaps in decisions and contacts.
Record the timings of each step against the RTO, every issue found, and who fixes it. Configuration drift is the usual enemy: a new volume outside the consistency group, a new started task missing from the DR automation, a certificate valid only at site A.
A service's agreed recovery time objective is 4 hours. In the DR test, systems took 2 hours to restart and applications 2.5 hours to validate. By how many minutes was the RTO missed?
Show a hint
Add the two durations and compare with 4 hours.
Show the solution
2 + 2.5 = 4.5 hours, which is 30 minutes over the 4-hour RTO.
Cost trade-offs
Each step toward zero RPO and zero RTO costs more: second and third sites, duplicate storage, network links, capacity that sits idle or lightly used, software licences, and skilled staff to run the automation. IBM and other vendors offer arrangements to reduce idle capacity costs, such as temporary capacity for DR use (IBM's Capacity Backup is one example), but terms depend on contract.
The architect's job is to put each service in the cheapest tier that meets its agreed targets, and to make the cost visible so the business decides knowingly. Complexity is a cost too: an elaborate design that nobody tests regularly can be less reliable than a simpler one that is exercised every quarter.
Modern relevance
Regulators in many countries now expect firms to prove operational resilience, including tested recovery of important business services within stated tolerances. That pushes DR from an infrastructure topic to a board-level one, and makes regular, evidenced testing a requirement rather than good practice.
Common mistakes
Without the scheduler's current plan, nobody knows which batch ran. Replicate it in the same consistency group as the data it drives.
Systems can IPL while applications still fail on missing certificates, hard-coded addresses or partner links. Include application validation in every test.
Some jobs already completed or transmitted. Check scheduler state, catalogs and partner confirmations before any rerun.
What you will see at work
- DR tests are scheduled events with timed runbooks, observers and a findings log tracked to closure.
- Batch designers are asked whether each job is restartable and what happens if it runs twice.
- Architects present availability tiers with costs so business owners sign off the RPO and RTO they are paying for.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.