Mainframe Path Start learning free
Applied11 min readLesson 3 of 3

Golden signals, alerts and end-to-end tracing

Good observability measures what users feel: latency, traffic, errors and saturation. Alerts should fire on symptoms that need a human, not on every threshold. Tracing and correlation IDs then show which tier, mainframe or not, is responsible.

The four golden signals

Google's site reliability engineering guidance popularised four golden signals for any service. They map well onto mainframe workloads.

SignalMeaningMainframe example
LatencyHow long requests take, including failed onesCICS transaction response time; DB2 elapsed time per thread
TrafficHow much demand there isTransactions per second; MQ messages put per minute
ErrorsHow many requests failTransaction abends, SQL errors, HTTP 500s from z/OS Connect
SaturationHow full the system isCPU busy, CICS max tasks reached, queue depth growing, WLM performance index above 1

Batch work needs a slightly different view: whether each job stream will meet its SLA, elapsed time against normal, and abends. Latency for batch is 'will the critical path finish on time'.

Designing alerts people trust

The fastest way to ruin on-call is an alert that fires all night and means nothing. People learn to ignore it, and then they ignore the real one.

Alert levels
Page now
Customer-facing service failingCritical path at risk of missing SLA
Ticket
Disk or log stream filling slowlyError rate above normal but stable
Dashboard only
CPU busy at expected peakSingle slow transaction

End-to-end transaction tracing

A distributed trace follows one request through many components. Each hop records a span with a start time, a duration and the same trace ID. Distributed systems pass the trace ID in HTTP headers; the W3C Trace Context standard defines a common header called traceparent.

One payment, many tiers
Mobile apptrace starts
API gatewaypasses trace ID
z/OS ConnectREST to CICS
CICS programtransaction
DB2SQL

Inside the mainframe, traditional tools record their own identifiers, such as the CICS task number and the DB2 correlation and accounting tokens. APM products with z/OS support, such as IBM Instana with IBM Z APM Connect, Dynatrace and others, connect these to the distributed trace so a single view shows time spent in each tier. Exact coverage depends on the product, version and subsystem.

Even without a full APM product, you can get much of the value by agreeing a correlation ID: generate one at the edge, pass it through the API and MQ message headers, and have the CICS or batch program log it. Then you can search every log for the same ID.

Following one correlation ID (illustrative)
14:05:02.118Z  api-gw     corr=7f3a9c  POST /payments  start
14:05:02.131Z  zosconn    corr=7f3a9c  CICS PAY1 tran=PAYA
14:05:02.140Z  cics       corr=7f3a9c  PAYA task 51234 start
14:05:04.902Z  cics       corr=7f3a9c  PAYA task 51234 end  2.762s
14:05:04.915Z  api-gw     corr=7f3a9c  200  total 2.797s

Here the gateway and z/OS Connect added only milliseconds; nearly all the time was inside the CICS task. The next question is whether that time was CPU, DB2 waits or lock contention, which the CICS and DB2 monitors can answer.

TRY IT YOURSELF

What is the W3C Trace Context HTTP header that carries the trace ID?

Show a hint

One word, all lower case, starting with trace.

Show the solution

traceparent

Common mistakes

Paging on resource thresholds alone

High CPU at peak can be normal and healthy on z/OS. Page on user-facing symptoms and use resource metrics to explain them.

Alerts with no runbook

A page without a known action wastes time at 3 a.m. Link every paging alert to clear first steps or downgrade it.

Not passing an ID across the boundary

If the trace or correlation ID stops at the API gateway, the mainframe part becomes a black box. Carry it into CICS, MQ and logs.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
← Getting mainframe data into enterprise platformsBack to Mainframe observability