Other People's Computers

The on-call who broke prod trying to fix it.

A senior engineer gets paged at 2am for a slow database. The dashboards say everything is fine. He restarts the connection pool. Then he fails over. Three hours later, he finds out what was actually wrong.

?Audio · subscribe to listen when the episode drops

Episode summary

The dashboards said CPU was fine, IOPS looked normal, connections were well below cap. But the application was timing out. He tried the obvious thing: restart the connection pool. Then the less-obvious thing: failover to the replica.

Then he found out what was actually wrong, three hours after the page first fired. The episode is about what he should have done in the first ten minutes that would have ended this.

We talk about the difference between symptoms and causes, why "restart it" is a reasonable default that breaks against the wrong shape of problem, and the runbook he wrote afterwards that reduced his team's MTTR by 40 percent. Not because it told them what to do, but because it told them what not to do.

Show notes

Back to all episodes