A list of the year's largest public outages is interesting, but it may not describe the incidents that mattered to your company. An annual review is more useful when it starts with customer effect and the services your team actually uses.

Build the incident set

Collect incidents from your own incident tracker, application alerts, vendor notices, and support cases. Include near misses when a fallback prevented customer impact. Keep the vendor's timeline separate from your internal timeline so that each can be corrected independently.

For every incident, record:

  • affected customer workflow
  • start, detection, acknowledgement, mitigation, and recovery times
  • internal service owner and external dependency
  • affected region or component, when confirmed
  • customer and business impact
  • fallback used and whether it worked
  • follow-up actions and their current owners

Group by failure mode

Provider names are useful for ownership, but failure modes reveal repeated engineering problems. Useful groups include identity, network, cloud region, data store, communications, payment, CI/CD, and certificate failures.

Look for patterns such as:

  • several products relying on the same identity or cloud provider
  • alerts reaching the wrong team
  • a provider notice arriving after application impact
  • failover blocked by missing data, credentials, or capacity
  • customer communication waiting on an unclear approval

Use comparable measures

Count incidents only under a documented rule. Track detection and recovery measures from the same event definitions each year. If the underlying data or monitoring coverage changed, explain the change rather than presenting the numbers as a continuous trend.

Vendor status observations are helpful context, but they are not the same as your application's availability. A provider can report a partial issue that does not affect you, or your account can fail while the public status remains normal.

Choose a short action list

The review should end with a small number of funded changes. Examples include assigning an unowned dependency, moving alerts to the correct team, testing a regional recovery plan, maintaining emergency identity access, or replacing an untested fallback.

Assign one owner and a due date to each change. Carry unfinished work into the regular reliability review rather than waiting for the next annual report.

Use incident history and monthly reports as external context, then reconcile that context with your own incident records.