Golden signals, alerts and end-to-end tracing
Good observability measures what users feel: latency, traffic, errors and saturation. Alerts should fire on symptoms that need a human, not on every threshold. Tracing and correlation IDs then show which tier, mainframe or not, is responsible.
The four golden signals
Google's site reliability engineering guidance popularised four golden signals for any service. They map well onto mainframe workloads.
| Signal | Meaning | Mainframe example |
|---|---|---|
| Latency | How long requests take, including failed ones | CICS transaction response time; DB2 elapsed time per thread |
| Traffic | How much demand there is | Transactions per second; MQ messages put per minute |
| Errors | How many requests fail | Transaction abends, SQL errors, HTTP 500s from z/OS Connect |
| Saturation | How full the system is | CPU busy, CICS max tasks reached, queue depth growing, WLM performance index above 1 |
Batch work needs a slightly different view: whether each job stream will meet its SLA, elapsed time against normal, and abends. Latency for batch is 'will the critical path finish on time'.
Designing alerts people trust
The fastest way to ruin on-call is an alert that fires all night and means nothing. People learn to ignore it, and then they ignore the real one.
- Alert on symptoms first. 'Card authorisation latency above 500 ms for 5 minutes' matters to customers. 'CPU above 90%' may be normal at peak.
- Use duration and rate, not single samples. A spike of one interval is usually noise.
- Every alert needs an action. Link it to a runbook entry. If nobody would do anything, make it a dashboard item, not a page.
- Route to the owner. A DB2 lock alert should reach the DB2 or application team, not only operations.
- Review regularly. After each incident, ask which alert should have fired earlier and which ones were noise.
End-to-end transaction tracing
A distributed trace follows one request through many components. Each hop records a span with a start time, a duration and the same trace ID. Distributed systems pass the trace ID in HTTP headers; the W3C Trace Context standard defines a common header called traceparent.
Inside the mainframe, traditional tools record their own identifiers, such as the CICS task number and the DB2 correlation and accounting tokens. APM products with z/OS support, such as IBM Instana with IBM Z APM Connect, Dynatrace and others, connect these to the distributed trace so a single view shows time spent in each tier. Exact coverage depends on the product, version and subsystem.
Even without a full APM product, you can get much of the value by agreeing a correlation ID: generate one at the edge, pass it through the API and MQ message headers, and have the CICS or batch program log it. Then you can search every log for the same ID.
14:05:02.118Z api-gw corr=7f3a9c POST /payments start 14:05:02.131Z zosconn corr=7f3a9c CICS PAY1 tran=PAYA 14:05:02.140Z cics corr=7f3a9c PAYA task 51234 start 14:05:04.902Z cics corr=7f3a9c PAYA task 51234 end 2.762s 14:05:04.915Z api-gw corr=7f3a9c 200 total 2.797s
Here the gateway and z/OS Connect added only milliseconds; nearly all the time was inside the CICS task. The next question is whether that time was CPU, DB2 waits or lock contention, which the CICS and DB2 monitors can answer.
What is the W3C Trace Context HTTP header that carries the trace ID?
Show a hint
One word, all lower case, starting with trace.
Show the solution
traceparent
Common mistakes
High CPU at peak can be normal and healthy on z/OS. Page on user-facing symptoms and use resource metrics to explain them.
A page without a known action wastes time at 3 a.m. Link every paging alert to clear first steps or downgrade it.
If the trace or correlation ID stops at the API gateway, the mainframe part becomes a black box. Carry it into CICS, MQ and logs.
What you will see at work
- SRE and operations teams define service-level objectives on latency and errors that include mainframe tiers.
- Incident bridges are shorter when a shared trace shows which tier consumed the time.
- Application teams add correlation IDs to COBOL logging when building new APIs.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.