Problem management and blameless review
Incidents get service back; problem management stops them coming back. That needs a blameless post-incident review, an honest look at contributing factors rather than a single culprit, and actions that are actually tracked to completion.
From incident to problem
A problem is the underlying cause of one or more incidents. Problem management is the process that investigates it and gets it removed. The split matters because the people restoring service at 3am are not in a position to do careful analysis, and because the same fault often shows up as several unrelated-looking incidents.
A known error is a problem whose cause and workaround are understood but not yet fixed. Recording it, with the workaround, means the next on-call engineer restores service in minutes instead of rediscovering the answer. Many sites link known errors directly from runbooks or from automation rules.
The post-incident review
For significant incidents (usually Sev 1 and Sev 2, and any near miss worth learning from) the team holds a post-incident review, sometimes called a post-mortem or retrospective. It is a structured meeting with a written output, held within days while memories are fresh.
- Build the timeline from the ticket, job logs, SYSLOG or OPERLOG, monitoring and chat history.
- Describe the impact in business terms: who was affected, for how long, what was lost or delayed.
- Ask how, not who: how did the system allow this, how was it detected, how was it restored.
- Identify contributing factors across technology, process and people.
- Agree actions with an owner and a date each, and record them in the problem record.
- Share the write-up so other teams learn from it.
Why blameless
A blameless review assumes people acted reasonably given what they knew at the time. That is not about being kind; it is about getting accurate information. If the operator who replied to the wrong console message expects to be punished, the review will hear a tidied-up story and the real weakness, perhaps a confusing message or a missing check in the runbook, will stay in place for the next person.
Root cause and contributing factors
The term root cause is useful but can mislead. Real mainframe incidents almost always have several contributing factors that only cause harm together. Stopping at one 'root cause' tends to produce a single fix and leave the other weaknesses in place.
| Factor | Example from a failed overnight posting run |
|---|---|
| Trigger | A new input file arrived with a record longer than the program expected |
| Technical | The program did not validate record length and abended with S0C7 on bad data |
| Detection | The monitor only alerted on job failure, not on the file arriving late and malformed |
| Process | The upstream change to the file layout was not shared with the receiving team |
| Recovery | The runbook for the job was out of date, so the restart took 90 minutes longer |
Techniques such as five whys and fishbone (Ishikawa) diagrams help a team keep asking past the first answer. Five whys is fast but tends to follow one path; a fishbone deliberately looks across categories such as people, process, technology and environment. Use whichever helps the conversation; the technique matters less than refusing to stop early.
Gathering evidence on z/OS
Good reviews rest on evidence rather than memory. Typical sources include:
- Job output and the system messages in JESMSGLG and JESYSMSG.
- The system log (SYSLOG, or OPERLOG in a sysplex) for console messages around the time.
- SMF records for job timings, dataset activity and resource use; RMF for performance.
- Subsystem logs and statistics from CICS, DB2, IMS or MQ.
- Scheduler history showing what ran, in what order, and what was held or forced.
- The change calendar: what changed in the hours or days before.
Actions that actually happen
The most common failure of problem management is not poor analysis but actions that are never completed. Good actions are specific ('add a record-length check to PGMX and a late-file alert to the scheduler'), owned by a named person, dated and tracked. Vague actions such as 'be more careful' or 'improve monitoring' are a sign the review stopped too early.
Common mistakes
'Operator error' explains nothing and teaches nothing. Ask what made the error easy to make and hard to catch, and fix that.
Most incidents need several factors to line up. Look across trigger, technical, detection, process and recovery before agreeing actions.
Actions without an owner and date quietly disappear, and the incident repeats. Record them in the problem record and review them until closed.
What you will see at work
- Junior engineers are often asked to build the timeline for a review from job logs, SYSLOG and the ticket, which is excellent training.
- Regulators and auditors in financial services may ask to see post-incident reports and evidence that actions were completed.
- Known-error records and their workarounds become a large part of what on-call staff use to restore service quickly.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.