Mainframe Path Start learning free
Core10 min readLesson 2 of 3

Problem management and blameless review

Incidents get service back; problem management stops them coming back. That needs a blameless post-incident review, an honest look at contributing factors rather than a single culprit, and actions that are actually tracked to completion.

From incident to problem

A problem is the underlying cause of one or more incidents. Problem management is the process that investigates it and gets it removed. The split matters because the people restoring service at 3am are not in a position to do careful analysis, and because the same fault often shows up as several unrelated-looking incidents.

How the records relate
Incidentssymptoms, restored
Problem recordone underlying cause
Known errorcause and workaround known
Changepermanent fix deployed

A known error is a problem whose cause and workaround are understood but not yet fixed. Recording it, with the workaround, means the next on-call engineer restores service in minutes instead of rediscovering the answer. Many sites link known errors directly from runbooks or from automation rules.

The post-incident review

For significant incidents (usually Sev 1 and Sev 2, and any near miss worth learning from) the team holds a post-incident review, sometimes called a post-mortem or retrospective. It is a structured meeting with a written output, held within days while memories are fresh.

  1. Build the timeline from the ticket, job logs, SYSLOG or OPERLOG, monitoring and chat history.
  2. Describe the impact in business terms: who was affected, for how long, what was lost or delayed.
  3. Ask how, not who: how did the system allow this, how was it detected, how was it restored.
  4. Identify contributing factors across technology, process and people.
  5. Agree actions with an owner and a date each, and record them in the problem record.
  6. Share the write-up so other teams learn from it.

Why blameless

A blameless review assumes people acted reasonably given what they knew at the time. That is not about being kind; it is about getting accurate information. If the operator who replied to the wrong console message expects to be punished, the review will hear a tidied-up story and the real weakness, perhaps a confusing message or a missing check in the runbook, will stay in place for the next person.

Root cause and contributing factors

The term root cause is useful but can mislead. Real mainframe incidents almost always have several contributing factors that only cause harm together. Stopping at one 'root cause' tends to produce a single fix and leave the other weaknesses in place.

FactorExample from a failed overnight posting run
TriggerA new input file arrived with a record longer than the program expected
TechnicalThe program did not validate record length and abended with S0C7 on bad data
DetectionThe monitor only alerted on job failure, not on the file arriving late and malformed
ProcessThe upstream change to the file layout was not shared with the receiving team
RecoveryThe runbook for the job was out of date, so the restart took 90 minutes longer

Techniques such as five whys and fishbone (Ishikawa) diagrams help a team keep asking past the first answer. Five whys is fast but tends to follow one path; a fishbone deliberately looks across categories such as people, process, technology and environment. Use whichever helps the conversation; the technique matters less than refusing to stop early.

Gathering evidence on z/OS

Good reviews rest on evidence rather than memory. Typical sources include:

Actions that actually happen

The most common failure of problem management is not poor analysis but actions that are never completed. Good actions are specific ('add a record-length check to PGMX and a late-file alert to the scheduler'), owned by a named person, dated and tracked. Vague actions such as 'be more careful' or 'improve monitoring' are a sign the review stopped too early.

Common mistakes

Naming a person as the root cause

'Operator error' explains nothing and teaches nothing. Ask what made the error easy to make and hard to catch, and fix that.

Stopping at the first cause

Most incidents need several factors to line up. Look across trigger, technical, detection, process and recovery before agreeing actions.

Leaving actions untracked

Actions without an owner and date quietly disappear, and the incident repeats. Record them in the problem record and review them until closed.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Incidents, severity and restoring serviceChange management, runbooks and automation →