Mainframe Path Start learning free
CoreOperations31 min

Incident, problem and change management

How production teams handle outages and stop them recurring: incident lifecycle and severity, restoring service first, blameless post-incident reviews, root cause versus contributing factors, safe change with back-out plans and freezes, runbooks and operations automation.

What you will be able to do

Lessons

  1. Incidents, severity and restoring service
    An incident is anything that interrupts or degrades a service the business relies on.
  2. Problem management and blameless review
    Incidents get service back; problem management stops them coming back. That needs a blameless post-incident review, an honest look at contributing factors rather than a single…
  3. Change management, runbooks and automation
    Most incidents follow a change, so production changes are assessed for risk, approved, scheduled and given a back-out plan.

Track your progress and earn a certificate.
Free account, 15 quiz questions for this subject.

Start this subject