Incidents and post-mortems

Confirmed incidents open themselves; manual incidents, post-mortems, and action items keep the story in one place instead of scattered across chat.

Confirmed incidents

A quorum-confirmed failure opens an incident with its start at the confirmation second. Every event lands on the timeline: region votes, notifications, acknowledgment, recovery. The timeline is the record your status page and your SLA evidence read from.

Manual incidents

Not everything a customer notices is a probe failure. Incidents can be opened by hand, with or without a bound monitor, curated, and explicitly published to the status pages you choose. Nothing publishes implicitly. Included from Sentinel, like incident command in general.

Post-mortems and action items

A post-mortem attaches to the incident: impact, root cause, lessons, plus action items with owners. It moves from draft to published deliberately, internal or public. The post-mortems policy states how we run our own.

Acknowledgment and ownership

Acknowledging takes the incident: escalation stops, other devices quiet down, reminders address the owner. The acknowledgment timestamp is part of the record, which is why the escalation ladder ends in an acknowledged incident, not in silence.