MN Michael Ngangom

Operational Excellence

From alert to evidence: incident operations by design

By Michael NgangomFeb 18, 20269 min read

Fast recovery and regulatory confidence depend on one coherent flow from signal through RCA and reporting.

During a payment incident, attention is the scarcest resource. The incident manager is coordinating recovery, translating technical signals into customer impact and watching a regulatory clock. Asking that person to recreate the same story in three systems is not governance. It is avoidable operational load.

01Signal
02Declare
03Stabilise
04Report
05Learn
One incident record should mature from an uncertain signal into evidence, reporting and learning.

Declaration must create structure, not paperwork

A useful declaration answers four things quickly: what service is affected, what customers may experience, who is in command, and when the next decision will be made. Precision improves as evidence arrives. Waiting for perfect certainty only delays coordination.

This thinking led directly to DoraIncident. I wanted declaration, ticket creation, a DORA countdown and the evolving incident timeline to be one connected action. The tool should carry the administrative weight while people carry the judgement.

Separate facts, hypotheses and decisions

Incident channels become noisy because observations and theories are mixed together. I use a simple discipline: timestamp facts, label hypotheses, and record decisions with an owner. That gives engineers room to investigate while keeping the command view trustworthy.

The same record can then support internal updates and regulatory reporting. DORA does not make good incident practice obsolete; it makes evidence quality and timing visible. If reporting is bolted on at the end, teams spend the final hours reconstructing what they already knew.

The review begins before recovery

A post-incident review should not be a hunt for the person who made the last change. It should explain why the system allowed that change to create the observed impact, why controls did or did not detect it, and which improvement will reduce repeat risk.

The most valuable outcome is not a polished RCA. It is a smaller gap between the next signal and the next correct decision.