Replication, backups, tape and GDPS
Data reaches the recovery site in two broad ways: continuous disk replication, synchronous or asynchronous, and point-in-time backups to disk or (virtual) tape. Automation such as GDPS ties them together so a whole site can be switched in a controlled way.
Getting data to the other site
Mainframe DR almost always depends on storage-based replication: the disk subsystems copy every write to a partner subsystem at another site, below the level of z/OS, DB2 or CICS. Because it works on disk volumes, everything on them is protected together: datasets, catalogs, DB2 logs, system libraries.
Synchronous replication is limited by physics: light in fibre takes time, and every write waits for the round trip. That is why synchronous partners are usually in the same metropolitan area, which in turn means a regional disaster could affect both. Many large sites therefore run three sites: two close together with synchronous replication, plus a distant one fed asynchronously.
Consistency is the whole point
A DB2 update writes to the log and later to the table space, often on different volumes. If the remote copy holds the table write but not the log write, the database at the recovery site is corrupt. Replication solutions therefore use consistency groups: sets of volumes whose remote copies are kept in the same write order, so that at any point the secondary looks like the primary would after a sudden power loss. DB2, IMS and CICS are designed to restart cleanly from that state, just as they would after a crash.
Point-in-time copies and backups
Alongside replication, sites keep point-in-time copies. Disk subsystems can take near-instant copies (for example IBM FlashCopy, with equivalents from other vendors), and some offer protected, immutable copies aimed at cyber recovery, such as IBM Safeguarded Copy. At the dataset level, DFSMSdss dumps and restores volumes and datasets, DFSMShsm manages backups and migration, and its ABARS function backs up groups of datasets for an application. DB2 and IMS have their own image copy and log-based recovery utilities, which give finer control than restoring whole volumes.
| Copy type | Protects against | Weakness |
|---|---|---|
| Synchronous replication | Site loss with no data loss | Copies corruption instantly; distance limited |
| Asynchronous replication | Regional site loss | Small data loss; still copies corruption |
| Point-in-time disk copy | Logical corruption, failed changes | Recovery point is when the copy was taken |
| Backup to tape or virtual tape | Long-term and offsite retention | Slower to restore; older recovery point |
| Database image copy plus logs | Damage to one database object | Needs DBA skill and good log retention |
Tape and virtual tape
Tape is still central to mainframe backup, but today it is mostly virtual tape: a virtual tape server presents tape drives to z/OS while storing the volumes on disk, and may move them later to physical tape or cloud object storage. Products include the IBM TS7700 family, Dell EMC DLm and Luminex. Virtual tape systems can usually replicate between sites themselves, so backup tapes created at the primary are already present at the recovery site.
Tape catalog and tape management data (for example in DFSMSrmm or Broadcom CA 1) must be recovered too. Without it, z/OS cannot tell which volume holds which dataset or whether it has expired.
GDPS: automation for the whole site
Replicating data is only half of DR. Someone has to stop the primary cleanly or detect it has gone, make the secondary disks usable, bring up systems in the right order and reconnect networks. GDPS is IBM's family of solutions, combining automation software with storage replication, that manages this for z/OS environments. Offerings cover synchronous metro configurations (GDPS Metro), long-distance asynchronous ones (GDPS Global), multi-site combinations, and continuous availability designs that run workloads actively in two sites.
At concept level, GDPS monitors the systems and replication, can freeze replication so the secondary stays consistent if a problem is detected, and runs scripted, tested procedures for planned site switches and unplanned failover. Other storage vendors provide their own automation and work with GDPS in some configurations; check what your site actually runs.
1. Confirm decision to invoke DR and record the time 2. Stop or confirm loss of production systems at site A 3. Make site B disks usable as primary 4. Start systems at site B in the agreed order 5. Start DB2, IMS, CICS and MQ; check restart messages 6. Switch network routing to site B 7. Hand over to application teams for validation
Common mistakes
Replication mirrors every change, including deletions and corruption. Keep separate point-in-time and protected copies for logical recovery.
Copies of a database's data and log taken at different moments can be unusable. Use consistency groups covering all related volumes.
Backup tapes are useless if the tape catalog and management data at the recovery site do not know what is on them.
What you will see at work
- Storage administrators own replication, point-in-time copies and virtual tape, and are central in every DR test.
- DBAs decide whether a damaged DB2 or IMS object is recovered from image copies and logs or from a disk-level copy.
- Architecture reviews ask which volumes are in which consistency group and whether new applications have been added to them.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.