Mainframe Path Start learning free
Applied6 min readLesson 5 of 5

Change, incidents and doing this sustainably

The last part of the job is the part that keeps you employed and rested: making changes safely, running incidents calmly, and steadily reducing the number of things that go wrong.

How a change normally moves

The promotion path
DevelopmentWrite and unit test
Test / QAFunctional and regression
UATBusiness acceptance
Change approvalCAB or standard change
ProductionDeploy, verify, monitor
Change typeMeans
StandardPre-approved, low risk, repeatable — raise and go
NormalAssessed and approved before a scheduled window
EmergencyFix now, approve retrospectively, with justification

What a good change record contains

Running an incident well

  1. State the impact in business terms, not technical ones: 'card authorisations are failing for 3% of customers', not 'PAYPOST1 abended S0C7'.
  2. Separate restoring service from finding the cause. Restore first where you safely can.
  3. Keep one person coordinating and communicating so the people investigating are not interrupted.
  4. Timestamp your actions as you go. Reconstructing them afterwards from memory is unreliable.
  5. Hold a blameless review afterwards and produce actions with owners and dates.

Reducing the work

Every recurring incident is a candidate for elimination. Keep a simple count of failures by job over a month, and the top three will be obvious. Typical fixes are unglamorous and highly effective:

Recurring problemDurable fix
Space abends on the same fileRight-size the allocation, or move to a managed class
The same bad record type every monthValidate and reject at the point of entry, with a clear message
A feed that arrives lateA dependency on arrival rather than a fixed time
Manual restart every MondayAutomate the restart, or fix the cause of the Sunday failure
Nobody knows how to fix XWrite the runbook the next time it happens

Common mistakes

An untested back-out plan

Verify the previous version still exists and can be restored before you need it at 2am.

Describing impact in technical terms

Business stakeholders need to know who is affected and how badly, not which job abended.

Skipping the review after a resolved incident

The fix that prevents recurrence usually comes from the review, not from the incident itself.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Rerun, restart, and not making it worseBack to Operations, scheduling and production support