Continuous availability: balancing, rolling IPLs, GDPS and common problems
Because work can run on any system, a Parallel Sysplex can spread load, survive the loss of one system and take systems down one at a time for maintenance. GDPS extends this across sites. Most sysplex problems are about structures, links or one sick system slowing the rest.
Spreading the work
Data sharing is only half the story. Something must also route each request to a system that has spare capacity. Several mechanisms do this, and most sites use more than one:
- WLM runs one service definition for the whole sysplex and gives routers recommendations about where capacity is available. See WLM.
- Sysplex Distributor in z/OS Communications Server presents one dynamic virtual IP address (DVIPA) and spreads incoming TCP connections across systems.
- VTAM generic resources let SNA users log on to one generic name, such as a CICS service, and be placed on any member.
- CICSPlex SM workload management routes transactions from routing regions to whichever target region is best placed.
- MQ shared queues use a pull model: any queue manager in the group can get the next message, so the least busy system naturally does more work.
- Db2 sysplex workload balancing spreads distributed (DDF) connections across the members of a data sharing group.
Rolling IPLs
Maintenance usually means a new set of system libraries and an IPL. In a sysplex the business does not see that IPL, because systems are restarted one at a time:
- Prepare and test the new system libraries (often a new system residence volume set) and any parmlib changes.
- Move work off the first system: stop new work arriving there and let routing send it to the others.
- Shut down subsystems on that system cleanly, then remove it from the sysplex with
VARY XCF,sysname,OFFLINE. - IPL it at the new level, restart its subsystems and check they rejoin their groups and structures.
- Watch it under load, then repeat for the next system.
During a rolling IPL, different levels run side by side, so IBM publishes coexistence requirements: fixes that the old level needs before it can share with the new one. Missing one is a classic cause of trouble. The SMP/E course covers how those fixes are found and installed.
Beyond one site: GDPS
GDPS is IBM's family of automation offerings for continuous availability and disaster recovery across sites. It combines disk replication with automation that manages systems, CFs and disks during planned and unplanned switches. One well-known feature, HyperSwap, switches systems to the secondary disks without an IPL when a primary disk subsystem fails. Offering names have changed over time; current ones include GDPS Metro, GDPS Global - GM and GDPS Continuous Availability. Other storage vendors provide compatible replication and their own automation, so check what your site actually runs.
| Replication style | Data loss after a disaster | Distance |
|---|---|---|
| Synchronous (for example Metro Mirror) | Effectively none for committed writes | Limited, typically metro distances |
| Asynchronous (for example Global Mirror) | Seconds of recent updates may be lost | Effectively unlimited |
Common sysplex problems
| Symptom | Likely causes | First checks |
|---|---|---|
| A system stops updating status and is partitioned out | Hang, spin loop, lost signalling paths | Console log, SFM policy action taken, dumps |
| Structure full or close to full | Growth, undersizing, a connector not cleaning up | D XCF,STR,STRNAME=name; RMF CF activity |
| Lost connectivity to a CF | Link failure or CF LPAR down | D CF and D XCF,CF; check structures rebuilt or switched to duplex copy |
| Everything slows when one system is sick | Sympathy sickness: the sick system holds locks or ENQs, or answers signals slowly | Find the holder; partition the sick system if needed |
| High lock contention | Hot data, or a lock structure too small causing false contention | Product statistics and RMF CF reports |
| Couple data set with no alternate | Alternate lost or never added | D XCF,COUPLE; add one with SETXCF COUPLE,ACOUPLE |
Parallel Sysplex remains the foundation for round-the-clock banking, card and payment systems. Newer ideas such as containers and APIs on z/OS sit on top of it rather than replacing it.
Common mistakes
If the cause is a shared structure or CF link, the slowdown follows the work to the other systems. Check D XCF,STR and RMF CF reports first.
Old and new levels must share the same structures and groups. Install the documented coexistence fixes on every system first.
It depends on disk replication, network, CF and system design and regular tested failovers. Untested DR plans fail on the day.
What you will see at work
- Change records for maintenance usually list a rolling IPL order and which workloads move during each step.
- On-call staff check D XCF and D CF displays early in any multi-system incident.
- DR tests, often GDPS-driven, run a few times a year and involve operations, storage, network and application teams.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.