Mainframe Path Start learning free
Applied11 min readLesson 3 of 3

Investigating performance problems

Good performance work follows a method: define the problem precisely, compare with a baseline, find where the time goes, change one thing, and measure again. Batch window problems are usually about the critical path and waiting, not raw CPU.

A structured method

Performance problems arrive vaguely: 'the system is slow'. The fastest route to a fix is a disciplined method that turns that into a measurable question.

  1. Define the symptom: which work, how slow, since when, compared with what? 'AUTH transaction 0.2s to 1.4s since 10:00' is a problem statement; 'CICS is slow' is not.
  2. Find the scope: one job, one service class, one LPAR, or everything? Wide scope suggests a shared resource or capping.
  3. Compare with a baseline: the same hour last week, or the last good run. What changed: volumes, code, configuration, hardware, schedule?
  4. Break down where the time goes: CPU, I/O, storage, locks, enqueues, waiting for another system. Use RMF Monitor III, Workload Activity, type 30 and subsystem accounting such as Db2 type 101.
  5. Form a hypothesis and test it with one change at a time, under change control.
  6. Measure again against the original symptom and record what you learned.
Narrowing the problem
Symptomwhat users or the schedule see
Scopewhich work, which system, when
BreakdownCPU, I/O, storage, locks, delays
Root causewhat changed and why it matters
Fix and verifyone change, measured

Batch window performance

Overnight, the question is whether the critical path finishes inside the batch window. Speeding up a job that is not on the critical path does not bring the end time forward.

Common causeTypical evidenceTypical remedy
Waiting for I/OElapsed time far above CPU, high EXCPBetter buffering, larger block sizes, fewer passes over the data
Serial dependenciesJobs wait in the schedule though resources are freeSplit work into parallel streams where data allows
ContentionENQ or Db2 lock waits, timeouts, deadlocksReschedule conflicting jobs, commit more often, review access order
Low WLM priorityCPU delay, batch class PI well above 1Review service class and importance through change control
Volume growthRun time rising steadily over monthsCapacity planning and algorithm changes, not just retries
CappingCPU delay with idle processors, R4HA at the limitReschedule heavy work, review caps

Anti-patterns

Common mistakes

Tuning the wrong job

In a batch window only the critical path sets the finish time. Check the schedule's dependencies before optimising anything.

Skipping 'what changed?'

Most sudden slowdowns follow a change: data volume, code, statistics, configuration or schedule. Look at change records early.

Declaring victory without measuring

Compare the next runs with the baseline. A fix that only helped once, or moved the delay elsewhere, is not a fix.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← MSU, capping and capacity planningBack to Performance and capacity management