Monitoring, observability and z/OS telemetry
Monitoring tells you when something you expected to break has broken. Observability is having enough data to work out why something unexpected is happening. On z/OS that data already exists in SMF, RMF, the system logs and product monitors; the job is knowing what each source tells you.
Monitoring versus observability
Monitoring is watching known things: is the CICS region up, is the batch stream on time, is CPU above a threshold. It answers questions you thought of in advance. Observability is the ability to ask new questions of the data when something odd happens, such as 'why are only card payments from one channel slow since 14:05?'. Monitoring is part of observability, not a different thing.
The industry usually talks about three kinds of telemetry. The mainframe has all three, and it had most of them long before the word observability was popular.
| Signal | What it is | Main z/OS sources |
|---|---|---|
| Metrics | Numbers measured over time: rates, response times, utilisation | RMF, SMF interval records, product monitors |
| Logs | Timestamped events and messages | SYSLOG and OPERLOG, job logs in JES, product logs, SMF event records |
| Traces | The path of one request through many components | Transaction monitors, CICS and DB2 trace facilities, APM tools with z/OS agents |
SMF: the system of record
SMF writes numbered record types. Different subsystems write different types, and many are only written if the site turns them on in the SMFPRMxx parmlib member. Records go to SMF datasets (the classic SYS1.MANx style) or, at most modern sites, to System Logger log streams.
| SMF type | Written by | Typical use |
|---|---|---|
| 30 | z/OS | Job and step resource use, completion codes |
| 70 to 79 | RMF | CPU, storage, I/O, workload activity (type 72 for WLM) |
| 80 | RACF | Security events such as violations |
| 100, 101, 102 | DB2 | Statistics, accounting per thread, performance trace |
| 110 | CICS | Performance per transaction, statistics |
| 115, 116 | IBM MQ | Statistics and accounting |
SMF is extremely detailed and trusted for capacity, chargeback and audit. Its weakness for observability has been delay: traditionally records were dumped and processed in batch hours later. z/OS now offers ways to read SMF data from in-memory resources close to real time, and products build on these, but what is available depends on z/OS level and site setup.
RMF: performance views
RMF has three monitors. Monitor I gathers long-term interval data and writes SMF 70 to 78 records for reports such as the Workload Activity report. Monitor II gives snapshot views of address spaces. Monitor III samples every second or so and shows short-term delays: who is waiting, and for what. RMF's Distributed Data Server can serve this data over HTTP, which is how many web tools and dashboards read it.
Logs: SYSLOG, OPERLOG and job output
Every system keeps a console log called SYSLOG. In a sysplex, OPERLOG merges the messages of all systems into one System Logger log stream, so you can see what happened across the sysplex in time order. Job output in JES (JESMSGLG, JESYSMSG, SYSOUT) is a log too, as are CICS and DB2 message logs. Messages have IDs, and the prefix tells you which component wrote them, which makes them good material for automated parsing.
Product monitors
Real-time monitors sit on top of these sources and add their own collectors. Common ones are IBM OMEGAMON (with agents for z/OS, CICS, DB2, IMS, MQ, networks and storage), BMC AMI Ops (formerly MainView) and Broadcom SYSVIEW. They show live response times, thread details and resource waits, raise alerts and often feed automation. Which one you use is a site decision; the concepts carry over.
Common mistakes
Many record types and subtypes are optional and controlled in SMFPRMxx and by each subsystem. Check what your site records before designing reports around it.
The problem may have started on another member. OPERLOG gives the merged, sysplex-wide view.
Observability is a property of how much useful data you have and can query. Monitors are one part of it.
What you will see at work
- Performance and capacity teams process SMF and RMF data daily for reports, chargeback and trend analysis.
- Operations and support teams watch product monitor consoles and get paged from their alerts.
- Security teams forward SMF type 80 and related records to the enterprise security platform.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.