Episode summary
The dashboards said CPU was fine, IOPS looked normal, connections were well below cap. But the application was timing out. He tried the obvious thing: restart the connection pool. Then the less-obvious thing: failover to the replica.
Then he found out what was actually wrong, three hours after the page first fired. The episode is about what he should have done in the first ten minutes that would have ended this.
We talk about the difference between symptoms and causes, why "restart it" is a reasonable default that breaks against the wrong shape of problem, and the runbook he wrote afterwards that reduced his team's MTTR by 40 percent. Not because it told them what to do, but because it told them what not to do.
Show notes
- The original incident timeline (sanitized).
- Three questions to ask before rolling back, restarting, or failing over.
- Why "the database is slow" almost never means the database is slow.
- A pattern for writing runbooks that do not get out of date.