CoreOperations31 min
Incident, problem and change management
How production teams handle outages and stop them recurring: incident lifecycle and severity, restoring service first, blameless post-incident reviews, root cause versus contributing factors, safe change with back-out plans and freezes, runbooks and operations automation.
What you will be able to do
- Run an incident from detection to closure, setting and revising severity by business impact
- Explain why restoring service comes before root cause, and capture diagnostics on the way
- Contribute to a blameless post-incident review that finds contributing factors and tracked actions
- Prepare a change with risk assessment, realistic back-out plan and verification, and know where automation fits
Lessons
- Incidents, severity and restoring serviceAn incident is anything that interrupts or degrades a service the business relies on.
- Problem management and blameless reviewIncidents get service back; problem management stops them coming back. That needs a blameless post-incident review, an honest look at contributing factors rather than a single…
- Change management, runbooks and automationMost incidents follow a change, so production changes are assessed for risk, approved, scheduled and given a back-out plan.
Track your progress and earn a certificate.
Free account, 15 quiz questions for this subject.