Incidents, severity and restoring service
An incident is anything that interrupts or degrades a service the business relies on. Good incident handling follows a known lifecycle, sizes the problem with a severity level, communicates on a rhythm, and puts restoring service ahead of finding the cause.
What counts as an incident
An incident is an unplanned interruption to a service, or a drop in its quality, that someone cares about. On the mainframe that might be a CICS region that stopped accepting transactions, a payment batch that abended on the critical path, a DB2 subsystem that is slow because of lock contention, or an MQ channel that stopped and left messages piling up. It does not need to be a full outage: card authorisations that take four seconds instead of a quarter of a second are an incident even though nothing is technically down.
Incident management exists because outages are inevitable on any platform, and the cost of an outage is dominated by how long it lasts and how chaotic the response is. A shared process means that at three in the morning nobody has to invent who decides, who talks to the business, and who is allowed to touch production.
The lifecycle
Most sites record incidents in an IT service management (ITSM) tool such as ServiceNow, BMC Helix or Jira Service Management. The ticket is the single record of what was seen, who was engaged, what was changed and when service came back. If it is not in the ticket, it did not happen as far as auditors are concerned.
Severity: sizing the problem
Every site defines its own severity (or priority) scale, usually from 1 (worst) to 4 or 5. Severity is driven by business impact and urgency, not by how technically interesting the fault is. It decides who gets paged, how often updates go out and whether a dedicated incident manager takes over.
| Severity | Typical meaning | Typical response |
|---|---|---|
| Sev 1 | Critical customer-facing service down or data at risk | Immediate page, incident manager, bridge call, updates every 30 minutes |
| Sev 2 | Major degradation or a critical deadline at risk | Page on-call, escalate if not improving, regular updates |
| Sev 3 | Limited impact, workaround exists | Handled in working hours or by on-call as time allows |
| Sev 4 | Minor, cosmetic or single user | Queue for normal work |
Restore service first
The single most important rule is that during an incident the goal is restoring service, not understanding the fault. Finding the root cause is valuable, but it belongs to problem management afterwards. If restarting a hung CICS region, rerouting work to another region in the sysplex, or backing out last night's change brings the service back, do that first, collect diagnostics on the way, and investigate later.
Restoring first does not mean destroying evidence. Before restarting something, capture what you can cheaply: the job log from SDSF, a system dump if the support team asks for one, the relevant console messages and the time. Many sites list the minimum diagnostics to capture in the runbook for each service.
On-call and communication
On-call engineers are the first technical responders. Good on-call practice is boring: acknowledge the page quickly, open or update the ticket, follow the runbook, and escalate early rather than late. Escalating is not failure; sitting on a Sev 1 alone for an hour is.
- One incident manager coordinates; engineers fix. Mixing the two roles slows both.
- Communicate on a rhythm, even with no news: what is affected, what is being done, when the next update is.
- Write in business terms for business readers: 'online banking payments are failing' beats 'CICS region PRDA1 is short on storage'.
- Keep a timeline in the ticket as you go; reconstructing it later is unreliable.
Measuring it
Sites track a handful of metrics. The best known is MTTR, usually read as mean time to restore (some sites mean repair or recovery; check the definition). Others include mean time to detect, mean time to acknowledge and the number of incidents by severity. Metrics are useful for trends; they become harmful when people start closing tickets early to make the numbers look good.
Common mistakes
Analysis can wait; users cannot. Restore service with the safest available action, capture diagnostics on the way, and hand root cause to problem management.
A fascinating abend in a test job is not a Sev 1. Severity follows business impact and urgency, using your site's definitions.
Stakeholders assume the worst when updates stop. Send short updates on the agreed interval, even if the update is 'still investigating, next update at 04:30'.
What you will see at work
- Operations and production support staff raise and work incident tickets daily, mostly at Sev 3 and 4, with occasional Sev 1 bridges.
- Many banks and insurers have a dedicated major-incident team that takes over coordination for Sev 1 and Sev 2 events.
- Interviewers frequently ask 'what do you do first when a critical job fails?' and expect to hear 'assess impact, restore service, communicate'.
Key terms
Check your understanding.
Take this lesson's quiz and save your progress. Free.