Cut incident rate 45% with SLOs + runbooks
The system was buckling under growth and the old design had run out of room — every fix bought weeks rather than quarters, and the next step up in traffic would have broken it outright. I owned the redesign end to end: the tradeoffs, the migration path, and a rollout that kept old and new live until the cutover was boring.
Added SLO-based alerting and rewrote the top runbooks; paging volume fell 65% in a quarter.