Cut model inference cost 20% with quantization
Search results were mediocre and clickthrough had been flat for a year, with nobody owning the number end to end. I took relevance from the ranking model through the evaluation set to the online test, and made every change defensible against a holdout instead of a hunch.
Quantized and batched serving; cost per 1k requests fell 25% at the same quality bar.