Mainframe Path Start learning free
Expert10 min readLesson 3 of 3

Consistency, testing and cost trade-offs

A DR copy is only useful if batch, online and system data line up and the team has proven they can recover from it. This lesson covers recovering in-flight batch and online work, how to test DR without risking production, and how to weigh availability against cost.

Consistency beyond the disk

A consistency group gives a crash-consistent copy: as if every system had stopped at one instant. Subsystems such as DB2, IMS, MQ and CICS then do what they do after any crash: read their logs, back out uncommitted units of work and complete committed ones. Online transactions in flight are simply backed out, and users retry. Batch is harder.

In-flight batch

A batch job half-way through updating a sequential file or a GDG has no automatic backout. After failover, operations must know exactly which jobs were running, which had completed, and which datasets they had created. That information lives in the scheduler's database, in the catalogs and in the job's own checkpoint records, so all of them must be in the replicated, consistent set.

Work typeState after failoverRecovery approach
Online DB2 / IMS / CICSUncommitted work backed out by subsystem restartAutomatic; users retry
Persistent MQ messagesRecovered from MQ logsAutomatic; check for duplicates if not under syncpoint
Batch updating databases with commitsCommitted to last commit pointRestart from checkpoint
Batch writing sequential files or GDGsPartial outputs may existDelete partial outputs, rerun step
File transfers to partnersUnknown whether deliveredConfirm with partner before resending

Designing batch to be idempotent or restartable at step level makes this manageable. Jobs that update in place without checkpoints, or that delete their input before the output is safe, turn a DR event into a data-repair exercise.

Testing DR

An untested DR plan is a hope. Tests should prove three things: the data copy is usable, the people and runbooks work, and the RTO is achievable. Common test styles:

Record the timings of each step against the RTO, every issue found, and who fixes it. Configuration drift is the usual enemy: a new volume outside the consistency group, a new started task missing from the DR automation, a certificate valid only at site A.

TRY IT YOURSELF

A service's agreed recovery time objective is 4 hours. In the DR test, systems took 2 hours to restart and applications 2.5 hours to validate. By how many minutes was the RTO missed?

Show a hint

Add the two durations and compare with 4 hours.

Show the solution

2 + 2.5 = 4.5 hours, which is 30 minutes over the 4-hour RTO.

Cost trade-offs

Each step toward zero RPO and zero RTO costs more: second and third sites, duplicate storage, network links, capacity that sits idle or lightly used, software licences, and skilled staff to run the automation. IBM and other vendors offer arrangements to reduce idle capacity costs, such as temporary capacity for DR use (IBM's Capacity Backup is one example), but terms depend on contract.

Cost and complexity rise with each tier
Backup restoreRPO hours, RTO days
Async mirroringRPO seconds-minutes, RTO hours
Sync + automationRPO near zero, RTO under an hour
Active/activeRPO seconds, RTO minutes

The architect's job is to put each service in the cheapest tier that meets its agreed targets, and to make the cost visible so the business decides knowingly. Complexity is a cost too: an elaborate design that nobody tests regularly can be less reliable than a simpler one that is exercised every quarter.

Modern relevance

Regulators in many countries now expect firms to prove operational resilience, including tested recovery of important business services within stated tolerances. That pushes DR from an infrastructure topic to a board-level one, and makes regular, evidenced testing a requirement rather than good practice.

Common mistakes

Leaving scheduler data out of replication

Without the scheduler's current plan, nobody knows which batch ran. Replicate it in the same consistency group as the data it drives.

Testing only the infrastructure

Systems can IPL while applications still fail on missing certificates, hard-coded addresses or partner links. Include application validation in every test.

Rerunning jobs blindly after failover

Some jobs already completed or transmitted. Check scheduler state, catalogs and partner confirmations before any rerun.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← DR topologies and replication choicesBack to High availability and DR architecture