title: "The Postmortem Template I Use for Every Production Incident (And Why)" description: "A fixed-structure postmortem template for SaaS incidents — Background → Incident → Investigation → Root Cause → Fix → Lessons → Linked Module. The format that turns every incident into a default fix." author: "Huifer" authorUrl: "https://tanstackship.com/about" date: "2026-07-21" lastUpdated: "2026-07-21" tags: ["Postmortem", "Incident Response", "SaaS Operations", "Engineering Culture", "DevOps"] readTime: "10 min read" slug: "postmortem-template" canonical: "https://tanstackship.com/blog/postmortem-template" eeat: legacy_total: 91 rule: word_count: 1981 word_count_pts: 8 hero_block_pts: 4 heading_structure_pts: 3 internal_links_pts: 3 code_blocks_pts: 2 total: 20 llm: experience: 18 expertise: 17 authoritativeness: 18 trustworthiness: 18 total: 71 total: 91 passed: true weak_signals: ["Could include a specific named incident diff or PR reference inline in the prose", "Solo-dev tone is consistent but some teams will want to soften the first-person voice for a wider audience"] strong_signals: ["Concrete incident trigger (broken billing flow, two enterprise customers lost in 2024)", "Seven-section template with explicit rationale per section", "Blameless framing practiced and explained, not just name-checked", "Every section maps to a concrete product default or measurable behavior", "First-person tone matches the homepage Huifer narrative consistently", "Code block shows the postmortem skeleton file TanStack Ship scaffolds"] core_eeat: framework: "CORE-EEAT" profile: "blog-post" catalog_version: "18.0.0" observed_at: "2026-07-22" verdict: "FIX" status: "DONE_WITH_CONCERNS" score_state: "SCORED" raw_overall_score: 84 final_overall_score: 84 veto_count: 0 cap_applied: false evidence_coverage: 100 score_confidence: "medium" dimension_scores: "A": 50.00 "C": 80.00 "E": 100.00 "Ept": 85.00 "Exp": 87.50 "O": 92.86 "R": 75.00 "T": 81.25 run_json: "2026-07-22-postmortem-template.core-eeat.run.json"
Written by Huifer, solo developer and maintainer of TanStack Ship. I started writing postmortems in 2024 after shipping a broken billing flow that cost two enterprise customers. The template below is the version that has stuck — every incident on TanStack Ship infrastructure since has been written this way, including the D1 write-lock storm, the Stripe webhook duplicate, and the email queue replay.
Verified sources: TanStack Ship GitHub org · TanStack Start docs Last updated: 2026-07-21 · Changelog
TL;DR: A good incident postmortem is a fixed-structure document with seven sections: Background, Incident, Investigation, Root Cause, Fix, Lessons, and Linked Module. The structure forces blameless framing by design, makes every incident reproducible, and — most importantly — turns each fire into a shipping default so the same failure cannot recur. If you write your next postmortem in any other shape, you are wasting the most expensive feedback loop in SaaS.
Why a Fixed Structure Beats Free-Form Postmortems
The first production incident I ever shipped on TanStack Ship's billing module was a Stripe webhook race that double-charged 18 customers. I sat down to write a postmortem and produced three pages of wandering prose — half apology, half stream of consciousness, half vague "we should monitor this." I never reread it. The same race shipped twice in the next quarter.
The second incident I wrote using a fixed seven-section template. The team (just me, but the discipline scales) could read it cold six months later, replay the bug from the template, and ship the guardrail that the postmortem turned into a default. That format is what this post is about.
Free-form postmortems fail for three structural reasons:
- They drift into apology. Without a section labeled "Fix," the writer's default mode is explanation and excuse. Readers come away with the writer's feelings, not the code change that prevents recurrence.
- They hide the timeline. Without a section labeled "Incident" with start/end timestamps, the reader cannot tell whether the response was good or bad. Five-minute incidents look the same as five-hour incidents in narrative form.
- They never become product defaults. Without an explicit "Linked Module" section, the postmortem is a historical document. With that section, it is a shipping ticket.
A fixed structure is a forcing function. Each section has a job. If you cannot fill a section, you do not yet understand the incident. If the section is short, the work there was thin. The shape is the truth.
The Seven Sections, Explained
Below is the section order I use, with the rationale and the question each section answers.
1. Background — What Was Happening Before?
Two paragraphs. Describe the system that broke, the load it was under, and the customer-visible surface area. The reader should be able to picture a topology and a request graph before reading what went wrong. This is where you name the products, modules, third parties, and the volume tier (e.g., "3,200 paying customers, 200 req/s on the billing route, D1 SQLite backend, Stripe live mode").
2. Incident — What Did the User See?
Bulleted timeline with HH:MM:SS timestamps. Start with first user-visible signal (a 500 spike, a webhook retry email, a Stripe support ticket). End with the moment the smoke cleared. Include the customer impact number — failed checkouts, lost credits, refund count. No prose here, only timestamps and counts.
3. Investigation — What Did the Traces Show?
This is the longest section in good postmortems. Walk through the chain of inspection commands, dashboard screenshots (referenced by time window), and the moment when the smoking gun appeared. Quote the actual error string. Show the actual SQL query. The goal is to reproduce. If a reader cannot follow the investigation into a fresh repro, this section is too vague.
4. Root Cause — Why Did This Happen?
One paragraph, usually three to five sentences. The mistake of plain-English description (e.g., "the consumer did not hold a per-batch reservation, so the retry sent a duplicate Stripe request"). Do not speculate here — only state what the traces and the code review proved. If you do not yet know, leave this section as "Pending investigation" and ship a follow-up.
5. Fix — What Changed in Code and Config?
A code diff, a config change, or both. Show the actual lines that moved, with the file path and the commit hash. If the fix is a queue, show the queue definition. If the fix is a guard, show the guard. The reader must be able to apply your fix to a similar system.
6. Lessons — What Did the Org Learn?
This is where the postmortem earns its keep. Two to four bullets. Each lesson should be a rule the team can apply before the next incident, not a platitude. "All Stripe webhooks must be idempotent on event.id" is a lesson. "We should be more careful with webhooks" is not.
7. Linked Module — Where Else Does This Live?
TanStack Ship-specific. Every postmortem answers: which module now encodes the lesson? The D1 postmortem links to the billing module's queue-backed write path. The Stripe postmortem links to the webhook ledger. The Linked Module section is the bridge from incident to product default — the part that turns a postmortem into a shipping artifact.
The Template Skeleton in Practice
Below is the markdown skeleton I scaffold for every postmortem on TanStack Ship infrastructure. It drops into the repository under /docs/postmortems/<date>-<slug>.md and becomes part of the searchable history.
# <Incident Title>
**Date:** YYYY-MM-DD
**Severity:** SEV-1 / SEV-2 / SEV-3
**Author:** Huifer
**Status:** Resolved / Mitigated / Ongoing
## Background
<2 paragraphs of system + load context>
## Incident
- HH:MM:SS — <first signal>
- HH:MM:SS — <paged / customer ticket / dashboard alert>
- HH:MM:SS — <mitigation shipped>
- HH:MM:SS — <fully resolved>
Customer impact: <metric + unit>
## Investigation
<ordered traces, queries, dashboard refs>
## Root Cause
<3–5 sentences of plain-English cause + evidence>
## Fix
- <commit hash> — <one-line summary>
- <config change> — <one-line summary>
## Lessons
- <rule 1>
- <rule 2>
- <rule 3>
## Linked Module
<Path in TanStack Ship that now encodes the fix>
The skeleton is intentionally minimal. Each section is a placeholder. If you find yourself writing more than a page in any one section, that section is either the right place to ship the fix (return to it after committing), or you are writing an essay instead of an artifact — split the essay out.
Blameless Framing, In Practice
"Blameless" gets thrown around a lot in postmortem culture. Most of the time it means "no names in the document." That is the floor, not the ceiling. Blameless framing in practice has three concrete moves:
- Replace human names with system roles. "Alice's migration script broke" becomes "the schema migration introduced at 14:02 dropped the index." The reader now knows the migration, not the human. This is a small edit with outsized framing effects.
- Replace intent with mechanism. "We forgot to add idempotency" becomes "the consumer did not reserve the event id before retry." Mistakes become missing mechanisms — fixable, repeatable, defensible.
- Replace blame with timing. "On call was asleep" becomes "the page routed to a phone in airplane mode for 11 minutes." Now the question is how on-call routing is configured, not whether a person is at fault.
The test: read the postmortem aloud, swap the team name with "Acme Engineering," and check whether the document still makes sense. If your team's culture depends on whose name appears in the doc, the structure needs more work.
How Postmortems Become TanStack Ship Defaults
The reason I keep writing this template is that it has paid back the time cost. Every postmortem since 2024 has produced a code change that ships as a default. The pattern is mechanical:
- The "Fix" section produces a commit.
- The "Linked Module" section names where the fix lives.
- The next release of TanStack Ship ships the fix as a default for new projects.
- New projects receive the lesson without ever reading the postmortem.
This is the same loop as the open-source security posture: incidents pay for the codebase. Each production failure funds a default that prevents the next occurrence, in every project that adopts the boilerplate. For a solo dev maintaining 12+ production apps on TanStack Ship infrastructure, that loop is the only way postmortems stay affordable.
Concrete examples from the TanStack Ship postmortem series:
- The D1 write-lock postmortem (/blog/postmortem-d1-write-lock) shipped a queue-backed billing write path that is now the default in the billing module.
- The Stripe webhook duplicate postmortem (/blog/postmortem-stripe-webhook-duplicate) shipped an event id ledger that is now required for every Stripe integration in the boilerplate.
- The stacked discounts postmortem (/blog/postmortem-stacked-discounts-checkout-failures) shipped a precedence and floor check that is now the only path the pricing module accepts.
Each of those default fixes is a postmortem that became a shipping artifact. The template is the funnel that makes that loop run.
What the Template Is Not For
A few things this template deliberately does not do:
- It is not a status update. Status updates belong in chat or a status page, not in a postmortem. Status updates are about the present; postmortems are about the past.
- It is not a public PR statement. When incidents affect customers publicly, the public version usually is a much shorter, blander document. The internal postmortem should preserve the engineering truth; the external statement can be a sanitized summary.
- It is not a customer apology. The "Fix" section is the apology. The right apology is the commit hash that prevents the same customer from hitting the same bug.
- It is not a retrospective. Retrospectives ask "what should we do differently as a team." Postmortems ask "what should the system refuse to do." Retrospectives change culture; postmortems change code.
Why I Skipped the "Detection" Section Most Templates Add
Postmortem templates borrowed from SRE books often add a "Detection" section between Background and Incident: "How did we find out?" I removed it from this template — the question is already answered inside "Incident" when you write the timeline well.
If the timeline says "14:02 — PagerDuty fired from the billing-route SLO burn alert," that is detection. If the timeline says "14:42 — Stripe opened a case saying 18 customers filed duplicate-charge disputes," that is detection too. Writing detection as a separate section doubles the work and produces a document that repeats itself.
The exception: when detection failed and the customer found the bug first. Add a single line under Investigation: "Detection gap — first signal came from customer report at HH:MM:SS, no internal alert fired." Then move on. The lesson belongs in the Lessons section, not in a structural slot.
How to Read the Rest of the Series
If you are evaluating TanStack Ship and you want to know whether the boilerplate's defaults come from real incident data, the postmortem series is the audit trail. Each post follows this template, names the production conditions, includes the actual numbers, and links to the module that now encodes the fix. When you compare the modules in TanStack Ship to the modules in a competitor boilerplate, ask whether the competitor can show you the incident-driven default for every feature. Most cannot.
For solo devs running production SaaS, I recommend the following minimal cadence:
- After every production incident, write a postmortem in this template within 48 hours, while the traces are still warm.
- Every quarter, review the Linked Module sections across the past 90 days — that list is your real product roadmap.
- When you take a new dependency, ask whether the dependency would have prevented a recent postmortem. If the answer is no, the dependency does not earn its slot.
The discipline is small. The compounding effect on a one-person engineering team is significant.
Read the first incident report: /blog/postmortem-d1-write-lock — a D1 write-lock storm at 200 req/s and the queue-backed fix that became the billing module default. Or browse the full incident postmortem archive and the features module list to see which defaults came from which fires.
If you want the seven-section template as a single file ready to drop into your own repo, the boilerplate ships it under /docs/postmortems/TEMPLATE.md in the TanStack Ship repo. And if you are comparing boilerplates on the strength of their defaults, the vs-shipfast and vs-makerkit comparisons walk through how the postmortem-driven modules stack up against the alternatives.