Designing for continuous availability
High availability means the service keeps running when a component fails or is changed. On z/OS that is done by duplicating every critical part, sharing data across systems in a Parallel Sysplex, spreading work so any system can take it, and applying maintenance one system at a time.
Availability is a design property
A single mainframe is very reliable hardware, but a service is a chain: network, LPAR, z/OS, subsystems, databases, disk, and the application itself. The service is only as available as the weakest link, and the most common outages are not hardware failures at all. They are planned changes, software defects and human error. Availability architecture therefore asks two questions of every component: what happens if it fails, and how do we change it without stopping the service.
A component whose failure stops the whole service is a single point of failure (SPOF). The architect's first job is to list them. Typical ones are surprisingly mundane: one CICS region that owns a file, one DB2 subsystem, one TCP/IP address that clients hard-code, one batch job that holds an exclusive enqueue during the online day.
Three layers of redundancy
Each layer depends on the one below. Cloning CICS regions is pointless if they all depend on one file that only one region can open. That is why data sharing is the heart of mainframe availability: it lets several z/OS images read and update the same data at the same time with full integrity.
How Parallel Sysplex data sharing works
In a Parallel Sysplex, up to 32 z/OS systems are joined through one or more coupling facilities: specialised LPARs that hold shared lock, cache and list structures. A DB2 data sharing group lets several DB2 members open the same databases; global locks and changed pages are coordinated through the coupling facility. IMS has its own data sharing, VSAM uses record-level sharing, and MQ can place queues in a queue-sharing group so any queue manager in the group can get the same message.
| Component | Sharing mechanism | What it removes |
|---|---|---|
| DB2 for z/OS | Data sharing group | Dependency on one DB2 subsystem |
| IMS | IMS data sharing and shared message queues | Dependency on one IMS control region |
| VSAM | RLS through SMSVSAM | One CICS region owning the file |
| IBM MQ | Shared queues in a queue-sharing group | Messages stranded on one queue manager |
| CICS | Cloned regions managed by CICSPlex SM | One region handling all transactions |
| TCP/IP | Dynamic VIPA and Sysplex Distributor | Clients tied to one system's address |
Workload balancing
Shared data only helps if work can land anywhere. Sysplex Distributor spreads incoming TCP/IP connections across systems, using WLM information about which systems have capacity. CICSPlex SM routes transactions to the healthiest target region. MQ shared queues let whichever server is free get the next message. When a system fails, new work simply stops being routed there, and the survivors pick it up.
Two things still need care: work that was in flight on the failed system, and retained locks. A failed DB2 member keeps locks on the rows it was updating until it is restarted (often on another system) and backs out or completes its units of work. Those rows are unavailable until then, which is why fast automated restart of a failed member matters.
Rolling maintenance
Most downtime used to be planned. With clones and shared data, maintenance can roll: drain work from one system, stop it, apply the change, bring it back, let it rejoin, then move to the next. The service never stops. This requires that adjacent levels of software can coexist in the same sysplex or data sharing group, which IBM documents per release as coexistence and fallback rules.
- Stop routing new work to system A (quiesce its CICS regions, remove it from distribution).
- Let in-flight work finish, then stop subsystems cleanly.
- Apply maintenance and IPL or restart.
- Verify, then re-enable routing to A.
- Repeat for system B, keeping enough capacity on the survivors throughout.
Modern relevance
The same principles drive cloud designs: redundancy, statelessness, automated routing. The mainframe difference is that shared data with full locking is built in, so even stateful, update-heavy workloads such as account balances can run active on several systems at once.
Common mistakes
Several CICS regions that all depend on one file-owning region or one non-shared DB2 still fail together. Trace every request to its data and make that layer shared.
Programs that rely on region-local state only run in one place. Find and remove affinities before claiming the workload is balanced.
If each system runs at high utilisation, losing one means the rest cannot absorb its work. Size so the remaining systems can carry the load during failures and rolling maintenance.
What you will see at work
- Availability reviews walk the request path from network to disk and mark each single point of failure with an owner and a fix.
- System programmers plan rolling maintenance windows system by system, checking coexistence rules for each product level.
- Developers are asked to prove new CICS programs have no affinities before they go into a cloned, workload-balanced region set.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.