AppliedChecklists
Overnight incident checklist
What to do, in order, when something fails at 3am.
First five minutes
- What failed, and what does it feed?
- Is it on the critical path? What is the SLA?
- Is there customer impact now, or only if it is not fixed by morning?
- Post a first update, even if it only says you are investigating.
Diagnosis
- Find the job on the spool; note the job ID.
- Read JESMSGLG and find the FIRST failing step.
- Read JESYSMSG for allocation, space and security messages.
- Read the program's own output for the record or key involved.
- Classify it: data, environment, resource, code, or upstream.
- Check the runbook for a known action.
- Ask what changed today.
Before acting
- Did the failed step update anything?
- Are the inputs still present and correct?
- Have any successors already run?
- Is a rerun safe, or is a restart from a step needed?
- State your plan to someone before executing it.
After
- Confirm the output is correct, not just that the job ended with RC=0.
- Release or rerun successors as needed.
- Update the ticket with what happened and what you did.
- Update the runbook while it is fresh.
- Raise the underlying defect if this will recur.
Stuck on something else?
Ask the community or search the full course.