Cut incident rate 60% with SLOs + runbooks
Two teams were blocked on the same shared bottleneck and each was waiting for the other to move first. I designed the contract between them and the migration path to reach it, so both could ship on their own schedule instead of renegotiating scope at every release.
Added SLO-based alerting and rewrote the top runbooks; paging volume fell 50% in a quarter.