Rickover · In development

Rehearse your incident response before the night it matters.

Rickover brings the teachable parts of naval-reactor operational discipline to your incident response: casualty procedures with a hard line between the actions you do from memory and the ones you do with the reference open, drills that run them under a clock against the real system, and a written after-action review every time. Built by a former naval nuclear reactor operator.

It's being built into a readiness service for your team — shared procedure libraries, scheduled drills, and qualification tracking, so you know your on-call rotation is ready before the incident, not after.

Why rehearsed response

Runbooks read fine at 2pm. They don't hold up at 3am.

Software operations has picked up incident commanders and blameless postmortems, but not the part of high-reliability operations that makes the biggest difference under pressure: a short, rehearsed sequence with a hard line between the actions you do from memory and the ones you do with the reference open — practiced under a clock until the response is fast and correct when it counts.

A reactor plant doesn't respond to a casualty by opening a wiki page. Rickover is that model, adapted for distributed systems: casualty procedures with a fixed, scannable structure, drills that run them for real, and a written after-action review every time.

What Rickover gives your team

From a written procedure to a rotation that has drilled it.

Casualty procedures

Failure responses in a fixed structure — immediate actions done from memory, supplementary actions done with the procedure open. Templates for the failure classes that recur across systems.

Timed drills

Walk the immediate actions on the real system with the clock running, optionally with the fault injected for real. Timing is compared to the procedure's expectations.

Blind drills

Withhold the immediate-action text so the operator states and performs each from memory, revealed and scored afterward. This is the real memory-items drill.

After-action reviews

Every drill produces a written critique comparing actual timing to the procedure's expectations, with a corrective-actions checklist — readiness you can point to.

Qualification tracking

Who has drilled which procedures, how recently, and how they did — so “qualified to stand this watch” is something you can check, with recency that decays.

Printable casualty cards

A one-page card per procedure, meant to be posted at the watch station — or wherever your team actually looks during an incident.

How it works

Make one runbook drillable. Not all of them.

You don't rewrite your runbooks. Rickover's on-ramp is one procedure at a time, starting with a failure that has actually paged you.

  • Pick a runbook for a failure class you already know recurs — a dependency brownout, a retry storm, a bad deploy.
  • Pull out the three-to-five actions where stopping to read makes it worse. Those become the immediate actions, memorised and drilled. The rest of the runbook becomes supplementary, verbatim.
  • Drill it. The after-action review tells you whether the split is right, and where the tooling or the procedure needs work.

Your services, thresholds, and flags fill into the templates from a single team config, so a procedure is yours in one step.

Shaped on real systems

Written and drilled against real incidents first.

Rickover's procedures are written and drilled against real client and research systems, not designed in the abstract and hoped onto your incidents afterward. The practice — the format, the drill, the after-action review — is in use now; the team readiness service is being built to fit how that actually goes.

Better with Meridian

Procedures that stay true as the system changes.

The place the reactor analogy strains is staleness: software changes every deploy, and a procedure written six months ago quietly stops matching reality. Rickover splits a procedure into a stable core — the symptoms and immediate actions you memorise and drill, which change only by a deliberate revision — and a living shell of supplementary actions, thresholds, and owning teams.

With Meridian connected, the living shell regenerates against your live architecture model, and a procedure that has drifted from the architecture it was written against is flagged. The memorised actions never change on their own. This integration is on the near-term roadmap.

Where it stands

In development, tested on real incidents as it's built.

  • Working now: the procedure format, the drill — open, blind, and with faults injected for real — and the after-action review, in use on real procedures.
  • Being built: the team readiness service — shared procedure libraries, drill scheduling, and a qualification model that tracks who has drilled what, how recently, and how they did.
  • On the roadmap: the Meridian integration for living procedures and drift detection.

It's not open to other teams' incidents yet. If you want to be an early user of the readiness service, or compare notes on rehearsed response, get in touch →

Origin

Grew out of a real question — and gets tested against real incidents.

Rickover started with a question worth answering for any team that would rely on it: does the operational discipline that keeps a reactor plant safe transfer to software? Using it on real incidents is how that question gets answered — and why it keeps improving instead of shipping once and going stale.

Read the research framing →  ·  Resolving Architecture home →