Mainframe Path Start learning free
Core10 min readLesson 1 of 3

Incidents, severity and restoring service

An incident is anything that interrupts or degrades a service the business relies on. Good incident handling follows a known lifecycle, sizes the problem with a severity level, communicates on a rhythm, and puts restoring service ahead of finding the cause.

What counts as an incident

An incident is an unplanned interruption to a service, or a drop in its quality, that someone cares about. On the mainframe that might be a CICS region that stopped accepting transactions, a payment batch that abended on the critical path, a DB2 subsystem that is slow because of lock contention, or an MQ channel that stopped and left messages piling up. It does not need to be a full outage: card authorisations that take four seconds instead of a quarter of a second are an incident even though nothing is technically down.

Incident management exists because outages are inevitable on any platform, and the cost of an outage is dominated by how long it lasts and how chaotic the response is. A shared process means that at three in the morning nobody has to invent who decides, who talks to the business, and who is allowed to touch production.

The lifecycle

A typical incident lifecycle (names vary by site and ITSM tool)
Detectalert, user, monitor
Log and classifyticket, severity
Respondtriage, escalate
Restoreservice back
Closeconfirm, record
Reviewproblem record

Most sites record incidents in an IT service management (ITSM) tool such as ServiceNow, BMC Helix or Jira Service Management. The ticket is the single record of what was seen, who was engaged, what was changed and when service came back. If it is not in the ticket, it did not happen as far as auditors are concerned.

Severity: sizing the problem

Every site defines its own severity (or priority) scale, usually from 1 (worst) to 4 or 5. Severity is driven by business impact and urgency, not by how technically interesting the fault is. It decides who gets paged, how often updates go out and whether a dedicated incident manager takes over.

SeverityTypical meaningTypical response
Sev 1Critical customer-facing service down or data at riskImmediate page, incident manager, bridge call, updates every 30 minutes
Sev 2Major degradation or a critical deadline at riskPage on-call, escalate if not improving, regular updates
Sev 3Limited impact, workaround existsHandled in working hours or by on-call as time allows
Sev 4Minor, cosmetic or single userQueue for normal work

Restore service first

The single most important rule is that during an incident the goal is restoring service, not understanding the fault. Finding the root cause is valuable, but it belongs to problem management afterwards. If restarting a hung CICS region, rerouting work to another region in the sysplex, or backing out last night's change brings the service back, do that first, collect diagnostics on the way, and investigate later.

Two different goals, two different timescales
Restore (incident)
Get users working againWorkarounds are fineMinutes matterOwner: on-call and incident manager
Fix (problem)
Stop it happening againNeeds analysis and testingDays or weeksOwner: problem manager and application team

Restoring first does not mean destroying evidence. Before restarting something, capture what you can cheaply: the job log from SDSF, a system dump if the support team asks for one, the relevant console messages and the time. Many sites list the minimum diagnostics to capture in the runbook for each service.

On-call and communication

On-call engineers are the first technical responders. Good on-call practice is boring: acknowledge the page quickly, open or update the ticket, follow the runbook, and escalate early rather than late. Escalating is not failure; sitting on a Sev 1 alone for an hour is.

Measuring it

Sites track a handful of metrics. The best known is MTTR, usually read as mean time to restore (some sites mean repair or recovery; check the definition). Others include mean time to detect, mean time to acknowledge and the number of incidents by severity. Metrics are useful for trends; they become harmful when people start closing tickets early to make the numbers look good.

Common mistakes

Hunting the root cause during a Sev 1

Analysis can wait; users cannot. Restore service with the safest available action, capture diagnostics on the way, and hand root cause to problem management.

Setting severity by technical interest

A fascinating abend in a test job is not a Sev 1. Severity follows business impact and urgency, using your site's definitions.

Going silent while you work

Stakeholders assume the worst when updates stop. Send short updates on the agreed interval, even if the update is 'still investigating, next update at 04:30'.

What you will see at work

Key terms

Check your understanding.
Take this lesson's quiz and save your progress. Free.

Take the lesson quiz
Problem management and blameless review →