Mainframe Path Start learning free
Applied10 min readLesson 1 of 3

Response time, throughput and bottlenecks

Performance means two things: how fast one piece of work finishes (response time) and how much work gets done per unit of time (throughput). Most problems come from a bottleneck in processor, I/O or storage, and response time grows sharply as a resource gets busy.

Two questions, two measures

Response time is how long one unit of work takes: a CICS transaction, a Db2 query, a batch job. Throughput is how many units complete per second, hour or night. They are related but not the same. A system can have excellent throughput while individual users wait, and fast response for a few users can collapse when volume doubles.

Response time is service plus waiting
Arrivesrequest or job
Queuewaits for a busy resource
Serviceuses CPU, does I/O
Completesresponse time = wait + service

Waiting is usually what changes. When a resource is lightly used, work rarely queues. As utilisation rises, queues grow faster and faster, so response time climbs steeply near saturation. This is why a disk that is 40% busy behaves very differently from one that is 85% busy.

The three classic bottlenecks

ResourceSymptomsWhere to look
Processor (CPU)High CPU delay for important work, PI above 1, latent demandRMF CPU Activity, Workload Activity, Monitor III delays
I/OLong DASD response times, jobs mostly in I/O wait, elapsed time far above CPU timeRMF DASD Activity, Monitor III device delays, type 30 EXCP counts
Storage (memory)Paging activity, storage delays, address spaces swapped or waiting for framesRMF Paging Activity, Monitor III storage delays

There are other kinds of waiting that are not about hardware: enqueue (ENQ) contention on datasets, Db2 lock waits, CICS waits for a busy region or a full queue. These show up as delay too, and buying a bigger machine does not fix them.

Reading I/O response time

RMF breaks DASD response time into parts, which tells you where the time goes:

Components of DASD response time in RMFWhat it means
IOSQ
Queued in z/OS waiting to start the I/O to the device
PEND
Started but waiting for a path or the device to accept it
DISC
Disconnect: the storage system is working on it, often a cache miss
CONN
Connect: data actually being transferred

High IOSQ suggests too much I/O aimed at one volume; high DISC often points to cache misses or back-end disk work in the storage system. Exact thresholds depend on the hardware, so compare with your own baseline.

Processor terms you will hear

A batch job's step summary showing I/O-bound behaviour (illustrative)
STEP     CPU-TIME   ELAPSED    EXCP
EXTRACT  00:00:48   00:31:12   4,882,016
SORT     00:01:10   00:03:05   210,544
LOAD     00:02:31   00:06:40   402,119
TRY IT YOURSELF

Which component of DASD response time measures the time spent queued in z/OS before the I/O is started? (four letters)

Show a hint

It combines 'I/O' and 'queue'.

Show the solution

IOSQ is the time I/O requests wait in z/OS before they can be started to the device.

Common mistakes

Treating high CPU busy as the problem

z/OS is meant to run busy. Look at whether important work is meeting its goals and what it is delayed by.

Tuning CPU when the job is waiting for I/O

If elapsed time is far larger than CPU time, the job is waiting. Find what it is waiting for before touching the code.

Buying capacity for a locking problem

Enqueue and lock contention are serialisation problems. More processors can even make them worse.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
MSU, capping and capacity planning →