Performance, monitoring and troubleshooting
A busy CICS region is limited by how many tasks it may run and how much storage it has. Learn the controls (MXT, transaction classes, storage limits), the signs of trouble (task queues, short-on-storage, runaway tasks) and where the evidence comes from.
The limits that shape a region
| Control | What it limits | Notes |
|---|---|---|
| MXT (SIT) / MAXTASKS (CEMT) | Total user tasks in the region | When reached, new tasks wait to be attached |
| Transaction classes | Active tasks for a group of transactions | MAXACTIVE caps them; PURGETHRESH purges excess queued tasks |
| DSALIM / EDSALIM | Storage for the dynamic storage areas below and above the 16 MB line | 64-bit storage is bounded by the region's MEMLIMIT |
| MAXOPENTCBS | Size of the L8/L9 open TCB pool | Since CICS TS 5.1 CICS calculates it from MXT (it can still be changed with CEMT SET DISPATCHER); too low and threadsafe or DB2 work waits for a TCB |
A transaction class (TRANCLASS) protects the region from a flood of one kind of work. If a batch-like enquiry starts 300 tasks at once, a class with MAXACTIVE(20) keeps the rest of the region responsive. Actual values depend entirely on site workload and are set by the CICS systems programmer.
Short-on-storage
When CICS cannot satisfy a storage request in one of its dynamic storage areas it goes short-on-storage (SOS). It stops attaching new tasks, tries to release cached storage, and writes storage manager (DFHSM) messages to the console. Response time collapses, and a long SOS can end with the region being recycled.
The cause is often not CICS itself: tasks wait on something else, accumulate, and consume storage. Raising MXT without fixing the wait simply lets the region fill faster.
Abends that point at performance
- AICA: a task ran too long without giving control back (a loop). The threshold is the ICVR system parameter, and the transaction definition can override it.
- AKCS: a task waited too long for a resource and was purged by its deadlock timeout (DTIMOUT on the transaction).
- Long DB2 waits or lock timeouts show up as SQL errors such as -911, handled in the program.
Where the evidence comes from
| Source | What it gives you |
|---|---|
| CEMT INQUIRE TASK / SYSTEM | Live view of tasks, their state and current limits |
| CICS statistics | Interval and end-of-day counts: times at MXT, SOS events, TCB use; formatted with DFHSTUP |
| CICS monitoring facility | Per-transaction CPU, response and wait times in SMF type 110 records |
| Trace and dumps | Internal and auxiliary trace (CETR to control) and transaction or system dumps for deep analysis |
| Monitors | IBM OMEGAMON for CICS, BMC AMI Ops (MainView) and Broadcom SYSVIEW give real-time views and alerts |
CEMT INQUIRE SYSTEM Maxtasks( 0250 ) Cicstslevel( 060100 ) Progautoinst( Autoinactive ) CEMT INQUIRE TASK Tas(0041207) Tra(PAY1) Fac(T123) Sus Ter Pri( 001 ) Tas(0041210) Tra(PAY1) Fac(T456) Sus Ter Pri( 001 ) ... 180 more PAY1 tasks suspended ...
Many suspended tasks for one transaction tell you where to look next: what are they waiting on? CEMT and monitors show the suspend reason, such as a DB2 thread, an enqueue or a remote region.
A troubleshooting order that works
- Confirm scope: one transaction, one region, or all regions in the CICSplex?
- Look at task counts and suspend reasons before changing anything.
- Check the console and job log for DFH messages, SOS and abends.
- Follow the wait downstream: DB2, MQ, remote region, network.
- Take dumps or trace if the cause is not clear, then act under the runbook.
Common mistakes
If tasks are waiting on something, more tasks means more storage used and a faster route to SOS. Find the wait first.
Purging clears the symptom and the evidence. Capture CEMT output or a dump first unless the runbook says otherwise.
Most CICS slowdowns are waits on DB2, MQ, another region or the network. Follow the suspend reason.
What you will see at work
- Production support watches task counts, MXT and SOS indicators on the monitor during peaks.
- Capacity reviews use SMF 110 data to see CPU and response time per transaction.
- Incident reports for online outages usually name the downstream wait, not just 'CICS was slow'.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.