Postmortems without blame
An incident review that hunts for a culprit teaches people to hide. One that hunts for causes teaches the system.
Why blame fails
When an incident review ends with a name, everyone learns the same lesson: do not be the person whose change was running when things broke. People become cautious about deploying, reluctant to report near misses, and careful to phrase their accounts defensively. The organization loses exactly the information it needs to prevent the next incident.
What a blameless review looks for
A blameless postmortem assumes that everyone involved acted reasonably given what they knew at the time. The question is not who made the mistake but why the mistake was easy to make and hard to catch.
Typical findings look like:
- The deploy tool allowed a change to production without a staging run.
- The alert fired, but it went to a channel nobody watches at night.
- The dashboard showed green because it measured the wrong thing.
- The runbook described a system that had changed six months earlier.
Each of these is fixable, and fixing them helps everyone.
A useful structure
- Summary: what happened and what the impact was, in plain language.
- Timeline: what was observed and done, with timestamps, written from the perspective of the people involved.
- Contributing factors: the conditions that allowed the incident, usually several rather than one.
- What went well: detection, communication, or recovery steps worth keeping.
- Action items: specific, owned, and dated.
Write the timeline carefully
The timeline is where blame most often sneaks back in. Describe actions and the information available when they were taken. "The on-call engineer restarted the service, which had resolved the same symptom twice before" explains a decision. "The engineer restarted the service without checking the logs" judges it.
Action items that matter
Favor changes to systems over changes to people. "Be more careful" is not an action item. "Add a check that blocks deploys while the error rate is elevated" is. Each item needs an owner and a date, and someone needs to follow up when the date arrives.
Share widely
Postmortems are most valuable when other teams read them. A failure in one service often reveals a pattern that exists elsewhere. Publish them internally, discuss the interesting ones in a regular review, and treat a well written postmortem as a contribution worth recognizing.
The long-term effect
Organizations that practice this consistently see more incidents reported, not fewer, at first. That is a good sign. It means problems are surfacing while they are still small.