Shipping software with a pilot's checklist
I love watching Air Crash Investigation, and I’d rather hear it in the Turkish dub than in the original, because of the voice of the Turkish dubbing artist Fatih Özacun. I don’t know where my love of aviation comes from, but I’ve always loved flying, aircraft, air travel, and idly watching a plane heading off into the distance.
This story could have been one of those episodes. On 30 October 1935, at Wright Field in Ohio, Boeing flew the most advanced bomber anyone had built. The Model 299 climbed, stalled and fell. Two of the five people on board died, the pilot among them.
The pilot was Major Ployer Hill, the Army Air Corps’ chief of flight testing. Few pilots anywhere had more experience. He had forgotten one small thing: the gust lock1 A gust lock pins the control surfaces so wind can’t slam them around while the aircraft is parked. Essential on the ramp, fatal on the runway and in the air. was still engaged. A newspaper called the 299 “too much airplane for one man to fly.”
More training was hard to justify; nobody was going to out-train the Army’s chief test pilot. So a group of test pilots wrote the steps for taxi, takeoff, flight and landing on a card. In Gawande’s account, the 299 then flew 1.8 million miles without an accident, and the Army ordered almost thirteen thousand of them. You know it as the B-17.
Two ways to forget
The 299 shows the first: too many steps for one person’s mind. A new system, or a migration on a table you’ve never touched. Nothing is routine yet, so everything has to be remembered at once.
The second is the opposite. The step you have done a thousand times is the one you skip, because skipping it has never hurt. You don’t test the rollback because it worked last time. You assume the flag is off because it usually is.
Both end with a step that only existed in someone’s head. Engineers need a better place to keep those steps than their own memory.
Two kinds of checklist
Daniel Boorman was designing flight-deck checklists at Boeing when Atul Gawande, a surgeon, went to see him while writing The Checklist Manifesto, a book about why checklists work in cockpits and how they could work in operating rooms. The first thing Boorman asks is which kind of checklist you need.
Read-do: read the line, then do it. This is for things you do rarely and under stress: the incident runbook, the database restore. You are not expected to remember; you are expected to follow.
Do-confirm: do the work from memory, then stop and check it against the list. This is the pull request template and the release check. It trusts your skill and catches the slip.
Software checklists often fail by being the wrong kind. A pull request template written as twelve numbered steps is a read-do list in a do-confirm moment: nobody opens a PR and follows a template line by line, so people tick the whole thing at the end without reading it. A restore runbook that says “verify the backups are healthy” is a do-confirm list in a read-do moment. At three in the morning you need the exact command, and the runbook only reminds you that one exists.
What makes one work
Boorman’s rules, as Gawande reports them:
- A pause point. A checklist belongs to a moment, like before merge or right after a deploy. Not “whenever”.
- Killer items. His term for the steps that hurt when skipped and get skipped anyway. Everything else goes somewhere else.
- Short. His rule of thumb is five to nine items. Spend much more than a minute at a pause point and people start shortcutting.
- Tested. Try it in a rehearsal, then revise it when real use exposes a gap.
“Short” sounds like it rules out a fifty-step database restore. It doesn’t. The checklist only has to be short enough to use at its pause point; the detailed steps live in the procedure it points to.
Asaf Degani and Earl Wiener, two human-factors researchers, studied airline checklists for NASA and added a rule about the answers: a response should give the item’s actual state or value, not the fact that you looked at it. They quote a pilot who answered “checked and set” during taxi and found out on the takeoff roll that the speed bugs were still on the approach setting. His verdict: “‘checked’ and ‘set’ can be said too easily without any sound verification.”
Let me unpack that. Before takeoff, the pilot sets the speeds from the takeoff card on little markers around the airspeed indicator: the bugs. The one that matters most is V2, the takeoff safety speed: the lowest speed at which the airplane can still climb safely if an engine fails. On takeoff the pilot flies to that marker. If the bug is still on the approach speed from the last landing, say 132, then the card says 145 but the pilot is flying to 132. The airplane climbs slower than it should, and if an engine quits, that gap matters exactly then.
Software needs one more rule on top of theirs: every item needs a clear condition for done. “Dashboards: watched” means nothing unless someone knows what to watch, for how long, and when to stop the rollout.
Here is a deploy card written with all of that in mind. The grey lines are checked by the pipeline; they stay on the card so the crew knows they exist. The rest belongs to a person. Tap a line to call it out.
Each response is a state someone can check. “Rollback-safe” means the previous release still runs against the new schema; if it doesn’t, the rollback line below it is a lie. “Last release” names the version you would go back to.
When the checklist becomes code
In software, some checklist items can become enforced checks in the delivery pipeline. Aviation does a version of this too. On the Boeing 777-9, the electronic checklist reads some items from the aircraft itself: if the folding wingtips aren’t extended, it marks the Before Takeoff checklist incomplete and raises a caution.
The clearest case I know is a gateway setup I built at work. Each backend team owns the files for its own routes, and those files compile, together with shared templates, into one configuration of more than 500 routes. Checking that the whole thing still builds after a one-line change is exactly the kind of step people skip. So on every pull request a pre-check builds the full configuration, and only a pull request that builds goes on to review.
What the pre-check can’t decide stays with a person: a code owner from the owning team has to approve before anything merges. The pipeline only checks the conditions someone wrote down. Whether the change is a good idea is still a human call, and that is where the attention should go.
A checklist needs an owner, like any other code. After an incident or a rehearsal, fix the lines that were unclear, and grey out the ones the pipeline now enforces so nobody checks them twice. The lines that need someone’s judgment stay, even when a script could tick them.
Gawande names the resistance plainly: using a checklist “feels beneath us,” an embarrassment for people who are supposed to be good.
So go back to that morning at Wright Field. In the cockpit sits the Army’s most experienced test pilot; nobody knows flying better. What causes the crash is a small lock left on the tail. The pilots who came after him weren’t more talented. They had a card, that’s all. With it, the same design flew 1.8 million miles without an accident and became the B-17.
Before you press deploy next time, ask yourself one thing: what’s my gust lock?