Triage: an abend at three in the morning
A method beats intuition here, especially when you are half awake. This sequence works on any job in any application, and it is what experienced people actually do.
The sequence
- Establish impact first. What does this job feed? Is it on the critical path? Is there a customer-facing consequence? This decides urgency and who you tell.
- Find the job on the spool. SDSF, filtered by job name. Note the job ID.
- Find the first failing step. Read JESMSGLG top down and stop at the first non-zero return code or abend. Everything after it is usually consequence.
- Read JESYSMSG. Allocation, security and space failures show up here with the dataset name.
- Read the program's own output. SYSOUT and SYSPRINT often name the record or key being processed when it failed.
- Classify the failure. Data, environment, code, resource, or upstream. This determines who fixes it and how quickly.
- Check the runbook. Many failures are known and have a documented action.
- Decide the action. Rerun, restart from a step, fix data and rerun, or escalate.
- Record what you did, in the ticket and in the runbook.
Classifying quickly
| Class | Typical signs | Usual owner |
|---|---|---|
| Data | S0C7, invalid record, a key not found, a count mismatch | You, often with the upstream team |
| Environment | Dataset not found, security violation, file held by CICS | You or the storage/security team |
| Resource | x37 space abend, out of region, spool full, extent limit | You, with storage |
| Code | A new failure right after a deployment | The development team |
| Upstream | A feed did not arrive, or arrived empty or malformed | The sending system's team |
A worked example
JESMSGLG IEF403I PAYPOST1 - STARTED - TIME=02.14.07 IEF142I PAYPOST1 STEP010 - STEP WAS EXECUTED - COND CODE 0000 IEF142I PAYPOST1 STEP020 - STEP WAS EXECUTED - COND CODE 0000 IEA995I SYMPTOM DUMP OUTPUT SYSTEM COMPLETION CODE=0C7 REASON CODE=00000007 TIME=02.31.44 SEQ=00841 CPU=0000 ASID=00A3 PSW AT TIME OF ERROR 078D1000 8A3C0F42 OFFSET=00000F42 IN EPNAME=PAYPOST IEF472I PAYPOST1 STEP030 - COMPLETION CODE - SYSTEM=0C7 SYSOUT (the program's own DISPLAYs) PAYPOST STARTED RUNDATE=20260930 RECORDS READ : 0000412 LAST KEY READ : 0044219877 Conclusion: step 030, S0C7, at offset F42, on record with key 0044219877. Next: look at that record's numeric fields in hex.
Communicating while you work
Give a first update quickly, even if it is 'investigating, no cause yet, impact appears limited to X'. Silence is read as absence. Then update on a rhythm — every 30 minutes on a major incident — with what you know, what you are doing, and when you will next update.
Common mistakes
Errors cascade. Find the first non-zero step; the rest is usually fallout.
'It is probably the same as last week' is right often enough to be dangerous. Confirm from the output.
If bad data can reach that program, it will again. Raise the underlying issue even when the immediate fix works.
Regular short updates keep everyone calm and keep you from being interrupted every five minutes.
What you will see at work
- Keep a personal log of every incident and its resolution. Within months it becomes the most useful document you own.
- The upstream team often knows about a bad feed before you do. Building that relationship saves hours.
- Escalation is not failure. Escalating early with a clear summary is a sign of competence.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.