Running and troubleshooting MQ
When messages stop flowing, the same few checks find most problems: is the queue growing, is anyone reading it, is the channel running, and has anything landed on the dead-letter queue?
The four questions
- Is the queue depth growing? Compare current depth with normal and with MAXDEPTH.
- Is anything reading it? Check how many handles have it open for input. Zero readers means the consumer is down or never started.
- Is the channel running? For remote destinations, check channel status and the transmission queue depth.
- Is anything on the dead-letter queue? Messages MQ could not deliver go there, each with a header explaining why.
Commands you will use
MQSC commands work through ISPF panels, the z/OS console with your queue manager's command prefix, or tools such as the MQ Explorer and web console. The prefix is set by each site, so a console command might look like +MQP1 DISPLAY QLOCAL(PAYMENTS.IN) CURDEPTH.
| Command | What it tells you |
|---|---|
| DISPLAY QLOCAL(name) CURDEPTH MAXDEPTH | How many messages are waiting and how many fit |
| DISPLAY QSTATUS(name) IPPROCS OPPROCS | How many handles have the queue open for input and output |
| DISPLAY CHSTATUS(name) | Whether a channel is RUNNING, RETRYING, STOPPED and so on |
| START CHANNEL(name) / STOP CHANNEL(name) | Start or stop a channel (usually by MQ administrators) |
| DISPLAY QLOCAL(*) WHERE(CURDEPTH GT 0) | Every local queue that currently has messages |
Poison messages
A poison message is one the reader cannot process: bad data makes it abend, the get is backed out, the message returns to the queue, and the reader fails on it again in a loop. The defence is the backout count in the MQMD: a well-written program checks it and, past the queue's BOTHRESH, moves the message to the backout queue (BOQNAME) and carries on. Without that logic one bad message can stop a whole stream.
The dead-letter queue
When MQ cannot deliver a message, for example because the target queue does not exist or is full, it puts the message on the dead-letter queue with a dead-letter header (MQDLH) that records the reason code and intended destination. Read the header first: the reason code usually points straight at the fix. Messages are then replayed with a dead-letter handler or a tool, according to site procedure.
Scenario: payments not arriving
Common mistakes
If a poison message caused the stop, the reader will fail again immediately. Look at why it stopped before restarting it.
The MQDLH reason code tells you why delivery failed. Replaying messages before fixing that cause just sends them back to the DLQ.
A deep queue with healthy readers may just be a peak. A queue with zero readers is an outage.
What you will see at work
- Monitoring tools alert on queue depth thresholds and channel status; production support responds using runbooks built on the checks in this lesson.
- Channel problems are often network or certificate issues: involve the network and security teams early.
- Every MQ incident review asks whether the application handles poison messages properly.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.