Mainframe Path Start learning free
Core11 min readLesson 3 of 3

Change management, runbooks and automation

Most incidents follow a change, so production changes are assessed for risk, approved, scheduled and given a back-out plan. Runbooks and operations automation then make routine responses consistent and fast.

Why change management exists

On a platform that runs payments, policies and payroll, a large share of incidents are triggered by something someone changed: a new program version, a JCL edit, a DB2 bind, a parameter in a system library, a new RACF rule. Change management does not exist to slow people down. It exists so that every production change is understood, reviewed by someone else, timed sensibly and reversible.

What a change record contains

ElementWhat it answers
Description and reasonWhat is changing and why
Risk and impact assessmentWhat could break, for whom, and how badly
Implementation planExact steps, in order, with who does each
Test evidenceHow it was proven outside production
Back-out planHow to return to the previous state, and how long that takes
VerificationHow you will know it worked in production
Schedule and approvalsWhen it runs and who signed off

Many sites classify changes as standard (pre-approved, low risk, repeatable, such as adding a user to an existing group), normal (assessed and approved, often by a change advisory board or CAB) and emergency (needed urgently to fix or prevent an incident, approved quickly and reviewed afterwards). Exact names and rules vary by site.

Back-out plans

A back-out plan is only real if it has been thought through for this specific change. 'Restore the old version' is not enough. Which load library member, which DB2 package version, which copy of the PARMLIB member? Does backing out the program also require backing out data the new version has already written?

Freeze windows and timing

A change freeze is a period when only emergency changes are allowed. Common freezes cover month end, quarter end, year end, peak trading days and major regulatory deadlines. Outside freezes, changes are scheduled in agreed windows, usually when online volumes are low and before the overnight batch, or with enough time afterwards to watch the next run.

Runbooks

A runbook is the written procedure for operating or recovering a specific thing: what an alert means, what to check, which actions are safe, who to call. Good runbooks are short, current, tested, and written for someone tired and unfamiliar with the application. They are also the starting point for automation: you cannot automate a response nobody has written down.

Operations automation

z/OS produces a constant stream of console messages. Operations automation products watch those messages and system events and act on them: replying to routine messages, restarting a failed started task, raising an alert to the monitoring tool, or running a scripted sequence to bring a subsystem down and back up in the right order. Widely used products include IBM Z System Automation, BMC AMI Ops Automation and Broadcom OPS/MVS, and many sites also use REXX-based rules. z/OS itself provides building blocks such as the message processing facility (MPF) for suppressing and flagging messages.

Message-driven automation at concept level
Eventconsole message, status change
Rulematch and conditions
Actionreply, restart, alert
Recordlog, ticket
An automation rule in plain words (illustrative, not real syntax)
WHEN   started task PAYAPI ends unexpectedly
AND    it has restarted fewer than 3 times in 1 hour
THEN   restart PAYAPI
       raise a Sev 3 alert with the end message text
ELSE   do not restart; raise a Sev 2 alert to on-call

Notice the limit on restarts. Automation that blindly restarts something every time it fails can hide a real problem or turn one failure into a loop. Automation rules are production code: they go through change management, are tested, and have owners.

Metrics that tie it together

Sites often track the change failure rate (changes that cause an incident or need backing out), the number of emergency changes, and MTTR. Rising emergency changes usually mean planning problems; a high change failure rate points at testing or review. These figures are signals for improvement, not targets for blame.

Common mistakes

Back-out plans that ignore data

Restoring a program does not undo records it has written. Plan for data, and decide in advance when you will fix forward instead.

Bundling many changes in one window

When something breaks you cannot tell which change did it. Keep windows focused and stagger risky changes.

Automating without limits

Unlimited automatic restarts can mask a fault or create a loop. Add thresholds, alerts and an owner, and test the rule like any other change.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Problem management and blameless reviewBack to Incident, problem and change management