Shipped a RAG assistant with 35% fewer hallucinations
Search results were mediocre and clickthrough had been flat for a year, with nobody owning the number end to end. I took relevance from the ranking model through the evaluation set to the online test, and made every change defensible against a holdout instead of a hunch.
Grounded answers in retrieval with citation checks; hallucination rate dropped 40% against the eval set.