Rerun, restart, and not making it worse
The riskiest moment in production support is the decision about what to run again. Rerunning something that already updated data can double-post transactions; skipping something that did not run leaves data missing. Both are worse than the original failure.
The three questions before any restart
- Did the failed step change anything? If it updated a file or a database and did not roll back, a plain rerun may double-apply.
- Are the inputs still there and still correct? A previous step may have deleted or moved them; a GDG may have rolled forward.
- Have successors already run? If a downstream job consumed partial output, fixing this job alone leaves inconsistency behind.
The options, and when each is right
| Action | Use when | Watch out for |
|---|---|---|
| Rerun from the top | Nothing was updated, or all updates rolled back | Steps that append rather than replace |
| Restart from a step | Earlier steps succeeded and their output still exists | Temporary datasets that were deleted at step end |
| Fix data, then rerun | A specific record caused the failure | Whether fixing data is permitted without a change record |
| Restore and rerun | Data is inconsistent and must be put back | Restore time; take a backup before you start |
| Skip and reconcile later | Non-critical, and a documented process exists | Never improvise this one |
What makes a job safely rerunnable
- Output is replaced, not appended. A delete step at the front, or a new GDG generation each run.
- Updates are checkpointed. The program records where it got to, in the same unit of work as the data.
- Processing is idempotent where possible. Applying the same transaction twice produces the same result as applying it once.
- Inputs are preserved. The job does not delete its own input until it has succeeded.
* resume at step 030, skipping the steps that already succeeded //PAYPOST1 JOB (ACCT01),'RESTART',CLASS=A,MSGCLASS=X, // RESTART=STEP030 //* //* WARNING: any &&TEMP datasets passed from earlier steps //* no longer exist. Steps 010-020 must be rerun, or the //* datasets recreated, before this will work.
Before you press anything
On anything touching financial or customer data, state your plan to someone else before executing it — even one sentence in a chat channel. 'Restarting PAYPOST1 from STEP030, step 020 output confirmed intact, no successors have run.' This takes fifteen seconds, catches the occasional serious mistake, and creates a record. It is the habit that separates people who are trusted with production from people who are not.
Common mistakes
Every rerun adds the data again. Check whether output is replaced or appended before doing anything.
Steps that consumed &&TEMP datasets will fail on allocation, or worse, find a stale permanent file.
A rerun may create a second new generation, so downstream jobs reading (0) get the wrong file.
State the plan to someone. Fifteen seconds, and it prevents the incidents people remember for years.
What you will see at work
- Each application should document its restart procedure per job. Where it is missing, writing it after an incident is the single most useful thing you can do.
- Some jobs have a dedicated recovery job that backs out a partial run. Find out which of yours do.
- Data fixes usually need a change record even in an emergency. Know the emergency process before you need it.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.