DR topologies and replication choices
Disaster recovery architecture decides how a second site takes over when the first is lost. The key choices are how the sites run (active/standby or active/active), how data gets there (synchronous, asynchronous or software replication) and how much automation drives the switch, all measured against RPO and RTO targets.
HA and DR are different problems
High availability handles the loss of a component inside a site. Disaster recovery handles the loss of the site itself: power, flood, fire, or a regional network failure. DR design starts from two business numbers. The recovery point objective (RPO) is how much recent data the business can afford to lose. The recovery time objective (RTO) is how long the service can be down. A payment system might demand RPO near zero and RTO in minutes; an internal reporting system might accept a day of each.
Topologies
In active/active designs each site runs its own sysplex and copy of the data, and a workload can be switched between them in seconds to minutes. The catch is that the application must tolerate two copies of the data being slightly out of step, and update conflicts must be designed out, often by deciding that each workload updates at only one site at a time.
Replication choices
| Method | How it works | Typical RPO | Distance | Trade-off |
|---|---|---|---|---|
| Synchronous disk mirroring | Write completes only when both copies have it | Near zero | Metro distances, limited by latency | Adds response time to every write |
| Asynchronous disk mirroring | Writes sent to the remote copy shortly after | Seconds to minutes | Any distance | Some data loss on disaster |
| Software replication | Database changes captured from logs and applied at the other site | Seconds, depending on lag | Any distance | Per-database setup; supports active/active |
| Tape or backup copies | Periodic backups shipped or replicated | Hours to a day | Any distance | Cheap, slow recovery |
Storage-based mirroring is offered by the main storage vendors: IBM (Metro Mirror and Global Mirror), Dell (SRDF) and Hitachi (TrueCopy and Universal Replicator), among others. Software replication for DB2 and IMS is offered by IBM (for example Q Replication for DB2 and IBM Data Replication for IMS) and other vendors. Many large sites combine methods: synchronous mirroring between two nearby sites for HA, plus asynchronous replication to a distant third site for regional disasters.
Consistency groups
Replicating every volume is not enough; the copy must be consistent. Databases rely on dependent writes: the log record is written before the data page it describes. If the remote copy captures the data page but not the log, DB2 cannot recover. A consistency group makes the replication solution freeze all related volumes at the same point so dependent-write order is preserved. Every volume a workload needs, including logs, catalogs and the scheduler's databases, must be in the same group.
Automation
A site failover involves hundreds of steps: stop replication, make the secondary disk usable, IPL systems, restart subsystems, redirect the network, start workloads. Doing this by hand under pressure is slow and error-prone. DR automation runs these as tested, scripted workflows. IBM's GDPS family is the best known on z/OS; at concept level it monitors the sites, manages the replication, and drives planned and unplanned site switches. Some variants can perform a HyperSwap, switching systems from primary to secondary disk without an outage when every system can reach both synchronously mirrored copies, whether the disks are in one site or across a stretched sysplex.
1 Detect loss of site A (heartbeat, storage alerts) 2 Freeze replication at last consistent point 3 Make site B disk read/write 4 IPL site B systems; restart DB2, IMS, CICS, MQ 5 Redirect network and DNS / dynamic VIPA 6 Resume batch from the scheduler's recovered state 7 Release online workloads to users
Choosing
Match the method to each workload rather than the whole estate. Card authorisation might justify active/active; the general ledger might use synchronous mirroring with automated restart; development systems might be restored from backups. Each tier costs differently, and that leads straight to the trade-offs in the next lesson.
Common mistakes
They are business decisions about loss and downtime. Get them agreed per service first, then design to them.
Logs, catalogs, scheduler databases, RACF and system volumes are also needed. Missing one makes the DR copy unusable.
Every write waits for the remote site, so latency grows with distance. Long distances need asynchronous methods and accept some data loss.
What you will see at work
- Architects build a table of services with agreed RPO and RTO, then map each to a replication tier and topology.
- Storage and system programmers manage which volumes sit in which consistency group and verify nothing has been added outside it.
- Operations teams own the site switch runbooks and the automation that executes them.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.