Investigating performance problems
Good performance work follows a method: define the problem precisely, compare with a baseline, find where the time goes, change one thing, and measure again. Batch window problems are usually about the critical path and waiting, not raw CPU.
A structured method
Performance problems arrive vaguely: 'the system is slow'. The fastest route to a fix is a disciplined method that turns that into a measurable question.
- Define the symptom: which work, how slow, since when, compared with what? 'AUTH transaction 0.2s to 1.4s since 10:00' is a problem statement; 'CICS is slow' is not.
- Find the scope: one job, one service class, one LPAR, or everything? Wide scope suggests a shared resource or capping.
- Compare with a baseline: the same hour last week, or the last good run. What changed: volumes, code, configuration, hardware, schedule?
- Break down where the time goes: CPU, I/O, storage, locks, enqueues, waiting for another system. Use RMF Monitor III, Workload Activity, type 30 and subsystem accounting such as Db2 type 101.
- Form a hypothesis and test it with one change at a time, under change control.
- Measure again against the original symptom and record what you learned.
Batch window performance
Overnight, the question is whether the critical path finishes inside the batch window. Speeding up a job that is not on the critical path does not bring the end time forward.
| Common cause | Typical evidence | Typical remedy |
|---|---|---|
| Waiting for I/O | Elapsed time far above CPU, high EXCP | Better buffering, larger block sizes, fewer passes over the data |
| Serial dependencies | Jobs wait in the schedule though resources are free | Split work into parallel streams where data allows |
| Contention | ENQ or Db2 lock waits, timeouts, deadlocks | Reschedule conflicting jobs, commit more often, review access order |
| Low WLM priority | CPU delay, batch class PI well above 1 | Review service class and importance through change control |
| Volume growth | Run time rising steadily over months | Capacity planning and algorithm changes, not just retries |
| Capping | CPU delay with idle processors, R4HA at the limit | Reschedule heavy work, review caps |
Anti-patterns
- Changing several things at once: if it improves, nobody knows why; if it gets worse, nobody knows which change to reverse.
- No baseline: without a normal day to compare with, every number looks suspicious or none does.
- Averages that hide peaks: a daily average CPU of 60% says nothing about the 10:00 peak.
- Making everything importance 1 in WLM, or putting work in SYSSTC: priorities stop meaning anything.
- Buying hardware for a software problem: lock contention, bad access paths and serial schedules are not fixed by capacity.
- Ignoring cost: a tuning change that saves elapsed time but raises the monthly R4HA peak may cost more than it saves.
- Fixing without recording: findings lost after the incident mean the same investigation happens again.
Common mistakes
In a batch window only the critical path sets the finish time. Check the schedule's dependencies before optimising anything.
Most sudden slowdowns follow a change: data volume, code, statistics, configuration or schedule. Look at change records early.
Compare the next runs with the baseline. A fix that only helped once, or moved the delay elsewhere, is not a fix.
What you will see at work
- Performance incidents are usually raised by operations when a critical job overruns or online response breaches its SLA.
- Investigations cross teams: operations, systems programmers, DBAs, storage and application developers each own part of the evidence.
- Regular performance reviews look at trends before they become overnight incidents.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.