Mainframe Path Start learning free
Applied11 min readLesson 1 of 3

Monitoring, observability and z/OS telemetry

Monitoring tells you when something you expected to break has broken. Observability is having enough data to work out why something unexpected is happening. On z/OS that data already exists in SMF, RMF, the system logs and product monitors; the job is knowing what each source tells you.

Monitoring versus observability

Monitoring is watching known things: is the CICS region up, is the batch stream on time, is CPU above a threshold. It answers questions you thought of in advance. Observability is the ability to ask new questions of the data when something odd happens, such as 'why are only card payments from one channel slow since 14:05?'. Monitoring is part of observability, not a different thing.

The industry usually talks about three kinds of telemetry. The mainframe has all three, and it had most of them long before the word observability was popular.

SignalWhat it isMain z/OS sources
MetricsNumbers measured over time: rates, response times, utilisationRMF, SMF interval records, product monitors
LogsTimestamped events and messagesSYSLOG and OPERLOG, job logs in JES, product logs, SMF event records
TracesThe path of one request through many componentsTransaction monitors, CICS and DB2 trace facilities, APM tools with z/OS agents

SMF: the system of record

SMF writes numbered record types. Different subsystems write different types, and many are only written if the site turns them on in the SMFPRMxx parmlib member. Records go to SMF datasets (the classic SYS1.MANx style) or, at most modern sites, to System Logger log streams.

SMF typeWritten byTypical use
30z/OSJob and step resource use, completion codes
70 to 79RMFCPU, storage, I/O, workload activity (type 72 for WLM)
80RACFSecurity events such as violations
100, 101, 102DB2Statistics, accounting per thread, performance trace
110CICSPerformance per transaction, statistics
115, 116IBM MQStatistics and accounting

SMF is extremely detailed and trusted for capacity, chargeback and audit. Its weakness for observability has been delay: traditionally records were dumped and processed in batch hours later. z/OS now offers ways to read SMF data from in-memory resources close to real time, and products build on these, but what is available depends on z/OS level and site setup.

RMF: performance views

RMF has three monitors. Monitor I gathers long-term interval data and writes SMF 70 to 78 records for reports such as the Workload Activity report. Monitor II gives snapshot views of address spaces. Monitor III samples every second or so and shows short-term delays: who is waiting, and for what. RMF's Distributed Data Server can serve this data over HTTP, which is how many web tools and dashboards read it.

Logs: SYSLOG, OPERLOG and job output

Every system keeps a console log called SYSLOG. In a sysplex, OPERLOG merges the messages of all systems into one System Logger log stream, so you can see what happened across the sysplex in time order. Job output in JES (JESMSGLG, JESYSMSG, SYSOUT) is a log too, as are CICS and DB2 message logs. Messages have IDs, and the prefix tells you which component wrote them, which makes them good material for automated parsing.

Product monitors

Real-time monitors sit on top of these sources and add their own collectors. Common ones are IBM OMEGAMON (with agents for z/OS, CICS, DB2, IMS, MQ, networks and storage), BMC AMI Ops (formerly MainView) and Broadcom SYSVIEW. They show live response times, thread details and resource waits, raise alerts and often feed automation. Which one you use is a site decision; the concepts carry over.

Where z/OS telemetry comes from
Dashboards, alerts, analyticsproduct consoles or enterprise platforms
Product monitorsOMEGAMON, AMI Ops, SYSVIEW
RMF and logsMonitor I, II, III; SYSLOG, OPERLOG, JES
SMF recordsz/OS, RMF, CICS, DB2, MQ, RACF
z/OS and subsystemswhere the work actually runs

Common mistakes

Assuming every SMF type is being written

Many record types and subtypes are optional and controlled in SMFPRMxx and by each subsystem. Check what your site records before designing reports around it.

Only reading one system's SYSLOG in a sysplex

The problem may have started on another member. OPERLOG gives the merged, sysplex-wide view.

Treating monitoring and observability as rival products

Observability is a property of how much useful data you have and can query. Monitors are one part of it.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
Getting mainframe data into enterprise platforms →