Mainframe Path Start learning free
Applied8 min readLesson 3 of 5

Triage: an abend at three in the morning

A method beats intuition here, especially when you are half awake. This sequence works on any job in any application, and it is what experienced people actually do.

The sequence

  1. Establish impact first. What does this job feed? Is it on the critical path? Is there a customer-facing consequence? This decides urgency and who you tell.
  2. Find the job on the spool. SDSF, filtered by job name. Note the job ID.
  3. Find the first failing step. Read JESMSGLG top down and stop at the first non-zero return code or abend. Everything after it is usually consequence.
  4. Read JESYSMSG. Allocation, security and space failures show up here with the dataset name.
  5. Read the program's own output. SYSOUT and SYSPRINT often name the record or key being processed when it failed.
  6. Classify the failure. Data, environment, code, resource, or upstream. This determines who fixes it and how quickly.
  7. Check the runbook. Many failures are known and have a documented action.
  8. Decide the action. Rerun, restart from a step, fix data and rerun, or escalate.
  9. Record what you did, in the ticket and in the runbook.

Classifying quickly

ClassTypical signsUsual owner
DataS0C7, invalid record, a key not found, a count mismatchYou, often with the upstream team
EnvironmentDataset not found, security violation, file held by CICSYou or the storage/security team
Resourcex37 space abend, out of region, spool full, extent limitYou, with storage
CodeA new failure right after a deploymentThe development team
UpstreamA feed did not arrive, or arrived empty or malformedThe sending system's team

A worked example

Reading the output of a real failure
JESMSGLG
IEF403I PAYPOST1 - STARTED - TIME=02.14.07
IEF142I PAYPOST1 STEP010 - STEP WAS EXECUTED - COND CODE 0000
IEF142I PAYPOST1 STEP020 - STEP WAS EXECUTED - COND CODE 0000
IEA995I SYMPTOM DUMP OUTPUT
  SYSTEM COMPLETION CODE=0C7  REASON CODE=00000007
  TIME=02.31.44 SEQ=00841 CPU=0000 ASID=00A3
  PSW AT TIME OF ERROR  078D1000  8A3C0F42
    OFFSET=00000F42 IN EPNAME=PAYPOST
IEF472I PAYPOST1 STEP030 - COMPLETION CODE - SYSTEM=0C7

SYSOUT (the program's own DISPLAYs)
PAYPOST STARTED  RUNDATE=20260930
RECORDS READ    : 0000412
LAST KEY READ   : 0044219877

Conclusion: step 030, S0C7, at offset F42, on record with key 0044219877.
Next: look at that record's numeric fields in hex.

Communicating while you work

Give a first update quickly, even if it is 'investigating, no cause yet, impact appears limited to X'. Silence is read as absence. Then update on a rhythm — every 30 minutes on a major incident — with what you know, what you are doing, and when you will next update.

Common mistakes

Reading the last error first

Errors cascade. Find the first non-zero step; the rest is usually fallout.

Guessing before reading

'It is probably the same as last week' is right often enough to be dangerous. Confirm from the output.

Fixing the data and moving on

If bad data can reach that program, it will again. Raise the underlying issue even when the immediate fix works.

Going quiet while investigating

Regular short updates keep everyone calm and keep you from being interrupted every five minutes.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← The batch schedule and the critical pathRerun, restart, and not making it worse →