Cut incident rate 55% with SLOs + runbooks
We were paging on-call most nights and the team had stopped believing the alerts, which meant real incidents got the same shrug as the noise. I made reliability the work instead of the afterthought — rebuilt the weakest paths, deleted the alerts that never mattered, and made the ones that stayed worth waking up for.
Added SLO-based alerting and rewrote the top runbooks; paging volume fell 55% in a quarter.