Applied6 min readLesson 5 of 5
Change, incidents and doing this sustainably
The last part of the job is the part that keeps you employed and rested: making changes safely, running incidents calmly, and steadily reducing the number of things that go wrong.
How a change normally moves
DevelopmentWrite and unit test
Test / QAFunctional and regression
UATBusiness acceptance
Change approvalCAB or standard change
ProductionDeploy, verify, monitor
| Change type | Means |
|---|---|
| Standard | Pre-approved, low risk, repeatable — raise and go |
| Normal | Assessed and approved before a scheduled window |
| Emergency | Fix now, approve retrospectively, with justification |
What a good change record contains
- What is changing and why, in language a non-specialist can follow.
- The exact objects affected: programs, JCL, procedures, packages, files.
- How it was tested, and what evidence exists.
- The back-out plan, tested where possible. This is the part most often skipped and most often needed.
- How you will verify success in production, and by when.
Running an incident well
- State the impact in business terms, not technical ones: 'card authorisations are failing for 3% of customers', not 'PAYPOST1 abended S0C7'.
- Separate restoring service from finding the cause. Restore first where you safely can.
- Keep one person coordinating and communicating so the people investigating are not interrupted.
- Timestamp your actions as you go. Reconstructing them afterwards from memory is unreliable.
- Hold a blameless review afterwards and produce actions with owners and dates.
Reducing the work
Every recurring incident is a candidate for elimination. Keep a simple count of failures by job over a month, and the top three will be obvious. Typical fixes are unglamorous and highly effective:
| Recurring problem | Durable fix |
|---|---|
| Space abends on the same file | Right-size the allocation, or move to a managed class |
| The same bad record type every month | Validate and reject at the point of entry, with a clear message |
| A feed that arrives late | A dependency on arrival rather than a fixed time |
| Manual restart every Monday | Automate the restart, or fix the cause of the Sunday failure |
| Nobody knows how to fix X | Write the runbook the next time it happens |
Common mistakes
An untested back-out plan
Verify the previous version still exists and can be restored before you need it at 2am.
Describing impact in technical terms
Business stakeholders need to know who is affected and how badly, not which job abended.
Skipping the review after a resolved incident
The fix that prevents recurrence usually comes from the review, not from the incident itself.
What you will see at work
- Change freezes around year end, month end and major business events are normal. Plan around them.
- Emergency changes are scrutinised afterwards. Document as you go so the review is easy.
- Bringing a monthly failure count with proposed fixes to your manager is one of the strongest things a new joiner can do.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.