Mainframe Path Start learning free
AppliedTroubleScheduler

Scheduled job failed overnight — scheduler failure workflow

Ended not OK, stuck waiting or running late: how to read the scheduler, find the cause, and choose rerun, force complete or hold safely.

What happened

You come in, or get paged, and the scheduler — Control-M, IBM Z Workload Scheduler or similar — shows a job in error, or a job that should have run is still waiting. Downstream work is blocked and the SLA clock is running.

Typical scheduler view (illustrative, product-neutral)
JOB        STATUS             DETAIL
PAYEXT01   Ended OK           RC 0000
PAYUPD02   Ended NOT OK       ABEND S0C7
PAYRPT03   Waiting            predecessor PAYUPD02
GLFEED04   Waiting            condition GL-FILE-READY not set
BANKSND05  Waiting            resource BANKLINE unavailable

What it means

The scheduler decides when jobs may start, based on time, predecessors, conditions and resources. It reports what z/OS told it about how the job ended. Each status points to a different problem, and the names differ by product.

SituationWhat it really means
Ended not OK / errorThe job ran and failed: abend, bad return code or JCL error
Waiting on predecessorAn upstream job has not completed successfully
Condition or resource not metA file, flag, or shared resource it depends on is not available
Late / at riskIt may still succeed, but will miss the deadline on the critical path

Typical causes

Symptoms

Where to look

How to diagnose

  1. Find the first failure in the chain; waiting jobs are usually victims, not causes.
  2. Read the job output to get the real error: which step, which abend or return code.
  3. Check what that step and earlier steps changed — files, Db2 tables, GDG generations.
  4. Check the runbook for restart points and whether the job is safe to rerun.
  5. For waits, see exactly which predecessor, condition or resource is missing.
  6. Estimate the impact on the critical path and deadlines.

How to fix

ActionSafe when
Rerun from the topThe job is idempotent, or its earlier updates are reversed first
Restart from a stepThe runbook names that step as a restart point and earlier steps completed correctly
Force complete / set OKThe work is genuinely not needed or was done another way, and downstream jobs will not process wrong or missing data
HoldYou need time to investigate and must stop dependants from starting

Fix the root cause first — the data, the program or the environment — then take the action, and watch the dependants start.

How to prevent

Production considerations

Interview question

A critical overnight job ended not OK and twenty jobs are waiting behind it. What do you do?

Stuck on something else?
Ask the community or search the full course.

Ask a question