Made experiment analysis self-serve for 10 teams
Search results were mediocre and clickthrough had been flat for a year, with nobody owning the number end to end. I took relevance from the ranking model through the evaluation set to the online test, and made every change defensible against a holdout instead of a hunch.
Built an A/B analysis layer; 4 product teams now read results without a data scientist in the loop.