Built a pipeline processing 10TB/day
We were flying blind on what the model did in production, with no idea it was wrong until a user complained loudly enough. I added monitoring on the predictions themselves, so drift and degradation showed up on a dashboard instead of in the support queue.
Rearchitected ingestion with Spark and dbt; the pipeline now processes 15TB/day with lineage.