Mainframe Path Start learning free
Applied11 min readLesson 3 of 3

Continuous availability: balancing, rolling IPLs, GDPS and common problems

Because work can run on any system, a Parallel Sysplex can spread load, survive the loss of one system and take systems down one at a time for maintenance. GDPS extends this across sites. Most sysplex problems are about structures, links or one sick system slowing the rest.

Spreading the work

Data sharing is only half the story. Something must also route each request to a system that has spare capacity. Several mechanisms do this, and most sites use more than one:

A typical online path through a Parallel Sysplex
Clientweb or mobile
Sysplex Distributorone DVIPA
CICS regionson SYSA, SYSB, SYSC
Db2 data sharingone member per system
Coupling facilitylocks and buffers

Rolling IPLs

Maintenance usually means a new set of system libraries and an IPL. In a sysplex the business does not see that IPL, because systems are restarted one at a time:

  1. Prepare and test the new system libraries (often a new system residence volume set) and any parmlib changes.
  2. Move work off the first system: stop new work arriving there and let routing send it to the others.
  3. Shut down subsystems on that system cleanly, then remove it from the sysplex with VARY XCF,sysname,OFFLINE.
  4. IPL it at the new level, restart its subsystems and check they rejoin their groups and structures.
  5. Watch it under load, then repeat for the next system.

During a rolling IPL, different levels run side by side, so IBM publishes coexistence requirements: fixes that the old level needs before it can share with the new one. Missing one is a classic cause of trouble. The SMP/E course covers how those fixes are found and installed.

Beyond one site: GDPS

GDPS is IBM's family of automation offerings for continuous availability and disaster recovery across sites. It combines disk replication with automation that manages systems, CFs and disks during planned and unplanned switches. One well-known feature, HyperSwap, switches systems to the secondary disks without an IPL when a primary disk subsystem fails. Offering names have changed over time; current ones include GDPS Metro, GDPS Global - GM and GDPS Continuous Availability. Other storage vendors provide compatible replication and their own automation, so check what your site actually runs.

Replication styleData loss after a disasterDistance
Synchronous (for example Metro Mirror)Effectively none for committed writesLimited, typically metro distances
Asynchronous (for example Global Mirror)Seconds of recent updates may be lostEffectively unlimited

Common sysplex problems

SymptomLikely causesFirst checks
A system stops updating status and is partitioned outHang, spin loop, lost signalling pathsConsole log, SFM policy action taken, dumps
Structure full or close to fullGrowth, undersizing, a connector not cleaning upD XCF,STR,STRNAME=name; RMF CF activity
Lost connectivity to a CFLink failure or CF LPAR downD CF and D XCF,CF; check structures rebuilt or switched to duplex copy
Everything slows when one system is sickSympathy sickness: the sick system holds locks or ENQs, or answers signals slowlyFind the holder; partition the sick system if needed
High lock contentionHot data, or a lock structure too small causing false contentionProduct statistics and RMF CF reports
Couple data set with no alternateAlternate lost or never addedD XCF,COUPLE; add one with SETXCF COUPLE,ACOUPLE

Parallel Sysplex remains the foundation for round-the-clock banking, card and payment systems. Newer ideas such as containers and APIs on z/OS sit on top of it rather than replacing it.

Common mistakes

IPLing a system to fix a sysplex-wide slowdown

If the cause is a shared structure or CF link, the slowdown follows the work to the other systems. Check D XCF,STR and RMF CF reports first.

Skipping coexistence fixes before a rolling IPL

Old and new levels must share the same structures and groups. Install the documented coexistence fixes on every system first.

Treating GDPS as a product you just switch on

It depends on disk replication, network, CF and system design and regular tested failovers. Untested DR plans fail on the day.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← The coupling facility and data sharingBack to Parallel Sysplex