Sample profile — not a real person

Grounded answers in retrieval with citation checks; hallucination rate dropped 40% against the eval set.

See the full story

Proof of work

2021

Shipped a RAG assistant with 35% fewer hallucinations

Search results were mediocre and clickthrough had been flat for a year, with nobody owning the number end to end. I took relevance from the ranking model through the evaluation set to the online test, and made every change defensible against a holdout instead of a hunch.

Grounded answers in retrieval with citation checks; hallucination rate dropped 40% against the eval set.

Machine LearningNLPPythonAirflow
2022

Added prediction monitoring across 10 models

The models were good in a notebook and useless in production, where nothing was reproducible and every handoff lost something. I built the path from experiment to serving that actually held up — versioned data, a real evaluation gate, and deploys that could be rolled back without ceremony.

Instrumented live predictions for drift and degradation on 7 models; problems surfaced on a dashboard, not in support.

Vector DatabasesSQL
2024

Automated retraining, keeping the model fresh

Everyone wanted 'AI in the product' and nobody had said what good would look like, so we were one demo away from shipping something confidently wrong. I set the evaluation first — the metric, the holdout, the bar to clear — so we could tell real progress from a nice-looking demo.

Scheduled retraining with drift detection and a promotion gate; the model stopped silently going stale between releases.

SQLSparkdbt
2022

Cut training cost 25% with spot + checkpointing

We were flying blind on what the model did in production, with no idea it was wrong until a user complained loudly enough. I added monitoring on the predictions themselves, so drift and degradation showed up on a dashboard instead of in the support queue.

Moved training to preemptible hardware with safe checkpointing; compute cost fell 55% at the same wall-clock.

SQLPythonAirflowPyTorch
2023

Cut dashboard data errors 55% with tests

Retraining was a manual ritual only one person knew, so the model quietly went stale between the times they had a free afternoon for it. I automated the loop — scheduled retraining, drift detection, and a gate that only promoted a model that actually beat the one already live.

Added dbt tests and freshness SLAs; data-quality incidents fell 45% in a quarter.

SQLPython

Sweat the boring details so the experience feels effortless.

Off the clock

Locked
PyTorchPandasMachine Learning
Locked
Python
Locked
LangChain
Locked
LangChain
Locked
Pythondbt

In public

SQLLLMsPython
Locked
Pandas
Locked

Practical bits

Timezones
Europe/Africa
Location
Germany
Languages
Greek, Vietnamese, Portuguese
Domains
Non-profit, IoT, Agriculture
Employment
Contract
Annual salary
Available with a company account
Social links
Available after unlocking

Reach out

Unlock with a company account to message directly and reveal:

  • Full name
  • Projects
  • In public
  • Social links

Find someone like this

Jump back to the board with this background already filtered in.