Mainframe Path Start learning free
Applied11 min readLesson 2 of 3

Replication, backups, tape and GDPS

Data reaches the recovery site in two broad ways: continuous disk replication, synchronous or asynchronous, and point-in-time backups to disk or (virtual) tape. Automation such as GDPS ties them together so a whole site can be switched in a controlled way.

Getting data to the other site

Mainframe DR almost always depends on storage-based replication: the disk subsystems copy every write to a partner subsystem at another site, below the level of z/OS, DB2 or CICS. Because it works on disk volumes, everything on them is protected together: datasets, catalogs, DB2 logs, system libraries.

Synchronous versus asynchronous replication
Synchronous
Write completes only after the remote copy confirmsRPO effectively zeroAdds latency to every writePractical over limited distances, typically metro rangeIBM: Metro Mirror; Dell: SRDF/S; Hitachi: TrueCopy
Asynchronous
Write completes locally; remote copy follows shortlyRPO of seconds to minutesLittle impact on application responseWorks over long distancesIBM: Global Mirror, z/OS Global Mirror; Dell: SRDF/A; Hitachi: Universal Replicator

Synchronous replication is limited by physics: light in fibre takes time, and every write waits for the round trip. That is why synchronous partners are usually in the same metropolitan area, which in turn means a regional disaster could affect both. Many large sites therefore run three sites: two close together with synchronous replication, plus a distant one fed asynchronously.

Consistency is the whole point

A DB2 update writes to the log and later to the table space, often on different volumes. If the remote copy holds the table write but not the log write, the database at the recovery site is corrupt. Replication solutions therefore use consistency groups: sets of volumes whose remote copies are kept in the same write order, so that at any point the secondary looks like the primary would after a sudden power loss. DB2, IMS and CICS are designed to restart cleanly from that state, just as they would after a crash.

Point-in-time copies and backups

Alongside replication, sites keep point-in-time copies. Disk subsystems can take near-instant copies (for example IBM FlashCopy, with equivalents from other vendors), and some offer protected, immutable copies aimed at cyber recovery, such as IBM Safeguarded Copy. At the dataset level, DFSMSdss dumps and restores volumes and datasets, DFSMShsm manages backups and migration, and its ABARS function backs up groups of datasets for an application. DB2 and IMS have their own image copy and log-based recovery utilities, which give finer control than restoring whole volumes.

Copy typeProtects againstWeakness
Synchronous replicationSite loss with no data lossCopies corruption instantly; distance limited
Asynchronous replicationRegional site lossSmall data loss; still copies corruption
Point-in-time disk copyLogical corruption, failed changesRecovery point is when the copy was taken
Backup to tape or virtual tapeLong-term and offsite retentionSlower to restore; older recovery point
Database image copy plus logsDamage to one database objectNeeds DBA skill and good log retention

Tape and virtual tape

Tape is still central to mainframe backup, but today it is mostly virtual tape: a virtual tape server presents tape drives to z/OS while storing the volumes on disk, and may move them later to physical tape or cloud object storage. Products include the IBM TS7700 family, Dell EMC DLm and Luminex. Virtual tape systems can usually replicate between sites themselves, so backup tapes created at the primary are already present at the recovery site.

Tape catalog and tape management data (for example in DFSMSrmm or Broadcom CA 1) must be recovered too. Without it, z/OS cannot tell which volume holds which dataset or whether it has expired.

GDPS: automation for the whole site

Replicating data is only half of DR. Someone has to stop the primary cleanly or detect it has gone, make the secondary disks usable, bring up systems in the right order and reconnect networks. GDPS is IBM's family of solutions, combining automation software with storage replication, that manages this for z/OS environments. Offerings cover synchronous metro configurations (GDPS Metro), long-distance asynchronous ones (GDPS Global), multi-site combinations, and continuous availability designs that run workloads actively in two sites.

At concept level, GDPS monitors the systems and replication, can freeze replication so the secondary stays consistent if a problem is detected, and runs scripted, tested procedures for planned site switches and unplanned failover. Other storage vendors provide their own automation and work with GDPS in some configurations; check what your site actually runs.

A site-switch script in plain words (illustrative, not GDPS syntax)
1. Confirm decision to invoke DR and record the time
2. Stop or confirm loss of production systems at site A
3. Make site B disks usable as primary
4. Start systems at site B in the agreed order
5. Start DB2, IMS, CICS and MQ; check restart messages
6. Switch network routing to site B
7. Hand over to application teams for validation

Common mistakes

Thinking replication is a backup

Replication mirrors every change, including deletions and corruption. Keep separate point-in-time and protected copies for logical recovery.

Replicating volumes without consistency

Copies of a database's data and log taken at different moments can be unusable. Use consistency groups covering all related volumes.

Forgetting tape management data

Backup tapes are useless if the tape catalog and management data at the recovery site do not know what is on them.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← RPO, RTO and business impactRecovering consistently and testing DR →