Mainframe Path Start learning free
Applied7 min readLesson 4 of 5

Rerun, restart, and not making it worse

The riskiest moment in production support is the decision about what to run again. Rerunning something that already updated data can double-post transactions; skipping something that did not run leaves data missing. Both are worse than the original failure.

The three questions before any restart

  1. Did the failed step change anything? If it updated a file or a database and did not roll back, a plain rerun may double-apply.
  2. Are the inputs still there and still correct? A previous step may have deleted or moved them; a GDG may have rolled forward.
  3. Have successors already run? If a downstream job consumed partial output, fixing this job alone leaves inconsistency behind.

The options, and when each is right

ActionUse whenWatch out for
Rerun from the topNothing was updated, or all updates rolled backSteps that append rather than replace
Restart from a stepEarlier steps succeeded and their output still existsTemporary datasets that were deleted at step end
Fix data, then rerunA specific record caused the failureWhether fixing data is permitted without a change record
Restore and rerunData is inconsistent and must be put backRestore time; take a backup before you start
Skip and reconcile laterNon-critical, and a documented process existsNever improvise this one

What makes a job safely rerunnable

The restart pattern in JCL
* resume at step 030, skipping the steps that already succeeded
//PAYPOST1 JOB (ACCT01),'RESTART',CLASS=A,MSGCLASS=X,
//            RESTART=STEP030
//*
//* WARNING: any &&TEMP datasets passed from earlier steps
//* no longer exist. Steps 010-020 must be rerun, or the
//* datasets recreated, before this will work.

Before you press anything

On anything touching financial or customer data, state your plan to someone else before executing it — even one sentence in a chat channel. 'Restarting PAYPOST1 from STEP030, step 020 output confirmed intact, no successors have run.' This takes fifteen seconds, catches the occasional serious mistake, and creates a record. It is the habit that separates people who are trusted with production from people who are not.

Common mistakes

Rerunning a job that appends

Every rerun adds the data again. Check whether output is replaced or appended before doing anything.

Restarting past a step whose temporary output is gone

Steps that consumed &&TEMP datasets will fail on allocation, or worse, find a stale permanent file.

Forgetting the GDG rolled forward

A rerun may create a second new generation, so downstream jobs reading (0) get the wrong file.

Acting alone at 3am on a financial job

State the plan to someone. Fifteen seconds, and it prevents the incidents people remember for years.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Triage: an abend at three in the morningChange, incidents and doing this sustainably →