Scheduled job failed overnight — scheduler failure workflow
Ended not OK, stuck waiting or running late: how to read the scheduler, find the cause, and choose rerun, force complete or hold safely.
What happened
You come in, or get paged, and the scheduler — Control-M, IBM Z Workload Scheduler or similar — shows a job in error, or a job that should have run is still waiting. Downstream work is blocked and the SLA clock is running.
JOB STATUS DETAIL PAYEXT01 Ended OK RC 0000 PAYUPD02 Ended NOT OK ABEND S0C7 PAYRPT03 Waiting predecessor PAYUPD02 GLFEED04 Waiting condition GL-FILE-READY not set BANKSND05 Waiting resource BANKLINE unavailable
What it means
The scheduler decides when jobs may start, based on time, predecessors, conditions and resources. It reports what z/OS told it about how the job ended. Each status points to a different problem, and the names differ by product.
| Situation | What it really means |
|---|---|
| Ended not OK / error | The job ran and failed: abend, bad return code or JCL error |
| Waiting on predecessor | An upstream job has not completed successfully |
| Condition or resource not met | A file, flag, or shared resource it depends on is not available |
| Late / at risk | It may still succeed, but will miss the deadline on the critical path |
Typical causes
- Program failure: an abend or a non-zero return code above the accepted limit.
- Bad or missing input: an empty, late or duplicate file from an upstream system.
- Environment: dataset contention, space, security or a region being down.
- Dependency not satisfied because an upstream job failed or was held.
- An external trigger (file arrival, condition from another system) never came.
- Volume growth making jobs run longer than the batch window.
Symptoms
- One job in error with a growing chain of waiting jobs behind it.
- Jobs waiting with no failures anywhere — usually a missing external condition or a held job.
- Everything succeeding but late — a performance or volume problem.
Where to look
- The scheduler's job details: status, return code, history of previous runs.
- The job output in SDSF or the archive: JESMSGLG, JESYSMSG, SYSOUT.
- The job's runbook for restart instructions and contacts.
- Upstream jobs and input files for missing or late data.
How to diagnose
- Find the first failure in the chain; waiting jobs are usually victims, not causes.
- Read the job output to get the real error: which step, which abend or return code.
- Check what that step and earlier steps changed — files, Db2 tables, GDG generations.
- Check the runbook for restart points and whether the job is safe to rerun.
- For waits, see exactly which predecessor, condition or resource is missing.
- Estimate the impact on the critical path and deadlines.
How to fix
| Action | Safe when |
|---|---|
| Rerun from the top | The job is idempotent, or its earlier updates are reversed first |
| Restart from a step | The runbook names that step as a restart point and earlier steps completed correctly |
| Force complete / set OK | The work is genuinely not needed or was done another way, and downstream jobs will not process wrong or missing data |
| Hold | You need time to investigate and must stop dependants from starting |
Fix the root cause first — the data, the program or the environment — then take the action, and watch the dependants start.
How to prevent
- Write a runbook for every critical job with restart points and contacts.
- Design jobs to be restartable, with checkpoints and idempotent updates.
- Validate input files early and fail fast.
- Monitor the critical path and alert on lateness, not just failures.
Production considerations
Interview question
A critical overnight job ended not OK and twenty jobs are waiting behind it. What do you do?
Stuck on something else?
Ask the community or search the full course.