Lessons From Our First Outage

By Jane Park · 2024-11-02 · 19 min read

performance python ai observability rust

This is an article about lessons-from-our-first-outage. Lessons From Our First Outage is a topic that comes up frequently in modern engineering practice. Practitioners often reach for it without first considering simpler alternatives. In this post, we examine when lessons-from-our-first-outage is the right call and when it isn't.

Why this matters

The literature on lessons-from-our-first-outage is vast and conflicting. Several industry surveys (summarized below) suggest that the average team over-uses lessons-from-our-first-outage by a factor of three to five. Our own experience matches this: across six projects, we found that lessons-from-our-first-outage delivered value in only the largest deployments.

The narrow case

lessons-from-our-first-outage wins when the workload matches the design. It loses when the workload is the wrong shape. The shape that fits is narrow but real. The shape that doesn't is broad and common.

Our recommendation

Start with the simplest tool that could possibly work. Resist the temptation to introduce lessons-from-our-first-outage until you can name the specific failure mode lessons-from-our-first-outage would prevent. When you do introduce it, make it observable. If you can't describe how it fails in two sentences, you are not ready to operate it.

Related reading

Categories

engineering architecture industry