Change management, runbooks and automation
Most incidents follow a change, so production changes are assessed for risk, approved, scheduled and given a back-out plan. Runbooks and operations automation then make routine responses consistent and fast.
Why change management exists
On a platform that runs payments, policies and payroll, a large share of incidents are triggered by something someone changed: a new program version, a JCL edit, a DB2 bind, a parameter in a system library, a new RACF rule. Change management does not exist to slow people down. It exists so that every production change is understood, reviewed by someone else, timed sensibly and reversible.
What a change record contains
| Element | What it answers |
|---|---|
| Description and reason | What is changing and why |
| Risk and impact assessment | What could break, for whom, and how badly |
| Implementation plan | Exact steps, in order, with who does each |
| Test evidence | How it was proven outside production |
| Back-out plan | How to return to the previous state, and how long that takes |
| Verification | How you will know it worked in production |
| Schedule and approvals | When it runs and who signed off |
Many sites classify changes as standard (pre-approved, low risk, repeatable, such as adding a user to an existing group), normal (assessed and approved, often by a change advisory board or CAB) and emergency (needed urgently to fix or prevent an incident, approved quickly and reviewed afterwards). Exact names and rules vary by site.
Back-out plans
A back-out plan is only real if it has been thought through for this specific change. 'Restore the old version' is not enough. Which load library member, which DB2 package version, which copy of the PARMLIB member? Does backing out the program also require backing out data the new version has already written?
Freeze windows and timing
A change freeze is a period when only emergency changes are allowed. Common freezes cover month end, quarter end, year end, peak trading days and major regulatory deadlines. Outside freezes, changes are scheduled in agreed windows, usually when online volumes are low and before the overnight batch, or with enough time afterwards to watch the next run.
- Avoid stacking unrelated changes in one window; if something breaks you will not know which one did it.
- Leave time to verify and, if needed, back out before the business day starts.
- Check the change calendar for other teams' changes to shared components such as DB2, CICS or MQ.
Runbooks
A runbook is the written procedure for operating or recovering a specific thing: what an alert means, what to check, which actions are safe, who to call. Good runbooks are short, current, tested, and written for someone tired and unfamiliar with the application. They are also the starting point for automation: you cannot automate a response nobody has written down.
Operations automation
z/OS produces a constant stream of console messages. Operations automation products watch those messages and system events and act on them: replying to routine messages, restarting a failed started task, raising an alert to the monitoring tool, or running a scripted sequence to bring a subsystem down and back up in the right order. Widely used products include IBM Z System Automation, BMC AMI Ops Automation and Broadcom OPS/MVS, and many sites also use REXX-based rules. z/OS itself provides building blocks such as the message processing facility (MPF) for suppressing and flagging messages.
WHEN started task PAYAPI ends unexpectedly
AND it has restarted fewer than 3 times in 1 hour
THEN restart PAYAPI
raise a Sev 3 alert with the end message text
ELSE do not restart; raise a Sev 2 alert to on-callNotice the limit on restarts. Automation that blindly restarts something every time it fails can hide a real problem or turn one failure into a loop. Automation rules are production code: they go through change management, are tested, and have owners.
Metrics that tie it together
Sites often track the change failure rate (changes that cause an incident or need backing out), the number of emergency changes, and MTTR. Rising emergency changes usually mean planning problems; a high change failure rate points at testing or review. These figures are signals for improvement, not targets for blame.
Common mistakes
Restoring a program does not undo records it has written. Plan for data, and decide in advance when you will fix forward instead.
When something breaks you cannot tell which change did it. Keep windows focused and stagger risky changes.
Unlimited automatic restarts can mask a fault or create a loop. Add thresholds, alerts and an owner, and test the rule like any other change.
What you will see at work
- Every production deployment, JCL change or parameter update you make will need a change record and usually a second person's review.
- Operations teams maintain automation rules for routine console messages and subsystem start-up and shut-down sequences.
- Month-end and year-end freezes shape release planning across the whole organisation.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.