Mainframe Path Start learning free
Expert11 min readLesson 1 of 3

Designing for continuous availability

High availability means the service keeps running when a component fails or is changed. On z/OS that is done by duplicating every critical part, sharing data across systems in a Parallel Sysplex, spreading work so any system can take it, and applying maintenance one system at a time.

Availability is a design property

A single mainframe is very reliable hardware, but a service is a chain: network, LPAR, z/OS, subsystems, databases, disk, and the application itself. The service is only as available as the weakest link, and the most common outages are not hardware failures at all. They are planned changes, software defects and human error. Availability architecture therefore asks two questions of every component: what happens if it fails, and how do we change it without stopping the service.

A component whose failure stops the whole service is a single point of failure (SPOF). The architect's first job is to list them. Typical ones are surprisingly mundane: one CICS region that owns a file, one DB2 subsystem, one TCP/IP address that clients hard-code, one batch job that holds an exclusive enqueue during the online day.

Three layers of redundancy

Removing single points of failure layer by layer
Workload routingSysplex Distributor, CICSPlex SM, MQ shared queues send work to any healthy system
Application and subsystem clonesSeveral CICS regions, DB2 members and IMS systems run the same work
Shared dataDB2 and IMS data sharing, VSAM RLS, coupling facility structures
Hardware and storageMultiple CPCs, duplexed coupling facilities, mirrored disk, redundant channels

Each layer depends on the one below. Cloning CICS regions is pointless if they all depend on one file that only one region can open. That is why data sharing is the heart of mainframe availability: it lets several z/OS images read and update the same data at the same time with full integrity.

How Parallel Sysplex data sharing works

In a Parallel Sysplex, up to 32 z/OS systems are joined through one or more coupling facilities: specialised LPARs that hold shared lock, cache and list structures. A DB2 data sharing group lets several DB2 members open the same databases; global locks and changed pages are coordinated through the coupling facility. IMS has its own data sharing, VSAM uses record-level sharing, and MQ can place queues in a queue-sharing group so any queue manager in the group can get the same message.

ComponentSharing mechanismWhat it removes
DB2 for z/OSData sharing groupDependency on one DB2 subsystem
IMSIMS data sharing and shared message queuesDependency on one IMS control region
VSAMRLS through SMSVSAMOne CICS region owning the file
IBM MQShared queues in a queue-sharing groupMessages stranded on one queue manager
CICSCloned regions managed by CICSPlex SMOne region handling all transactions
TCP/IPDynamic VIPA and Sysplex DistributorClients tied to one system's address

Workload balancing

Shared data only helps if work can land anywhere. Sysplex Distributor spreads incoming TCP/IP connections across systems, using WLM information about which systems have capacity. CICSPlex SM routes transactions to the healthiest target region. MQ shared queues let whichever server is free get the next message. When a system fails, new work simply stops being routed there, and the survivors pick it up.

Two things still need care: work that was in flight on the failed system, and retained locks. A failed DB2 member keeps locks on the rows it was updating until it is restarted (often on another system) and backs out or completes its units of work. Those rows are unavailable until then, which is why fast automated restart of a failed member matters.

Rolling maintenance

Most downtime used to be planned. With clones and shared data, maintenance can roll: drain work from one system, stop it, apply the change, bring it back, let it rejoin, then move to the next. The service never stops. This requires that adjacent levels of software can coexist in the same sysplex or data sharing group, which IBM documents per release as coexistence and fallback rules.

  1. Stop routing new work to system A (quiesce its CICS regions, remove it from distribution).
  2. Let in-flight work finish, then stop subsystems cleanly.
  3. Apply maintenance and IPL or restart.
  4. Verify, then re-enable routing to A.
  5. Repeat for system B, keeping enough capacity on the survivors throughout.

Modern relevance

The same principles drive cloud designs: redundancy, statelessness, automated routing. The mainframe difference is that shared data with full locking is built in, so even stateful, update-heavy workloads such as account balances can run active on several systems at once.

Common mistakes

Cloning regions but not the data path

Several CICS regions that all depend on one file-owning region or one non-shared DB2 still fail together. Trace every request to its data and make that layer shared.

Ignoring transaction affinities

Programs that rely on region-local state only run in one place. Find and remove affinities before claiming the workload is balanced.

Running survivors too hot

If each system runs at high utilisation, losing one means the rest cannot absorb its work. Size so the remaining systems can carry the load during failures and rolling maintenance.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
DR topologies and replication choices →