Sample profile — not a real person

Grounded answers in retrieval with citation checks; hallucination rate dropped 20% against the eval set.

See the full story

Proof of work

2026

Shipped a RAG assistant with 50% fewer hallucinations

Search results were mediocre and clickthrough had been flat for a year, with nobody owning the number end to end. I took relevance from the ranking model through the evaluation set to the online test, and made every change defensible against a holdout instead of a hunch.

Grounded answers in retrieval with citation checks; hallucination rate dropped 20% against the eval set.

PythonPandasGenerative AI
2023

Cut model inference cost 15% with quantization

We were flying blind on what the model did in production, with no idea it was wrong until a user complained loudly enough. I added monitoring on the predictions themselves, so drift and degradation showed up on a dashboard instead of in the support queue.

Quantized and batched serving; cost per 1k requests fell 30% at the same quality bar.

TensorFlowPython
2016

Cut dashboard data errors 65% with tests

Retraining was a manual ritual only one person knew, so the model quietly went stale between the times they had a free afternoon for it. I automated the loop — scheduled retraining, drift detection, and a gate that only promoted a model that actually beat the one already live.

Added dbt tests and freshness SLAs; data-quality incidents fell 25% in a quarter.

TensorFlowPyTorch
2019

Raised model accuracy 16%→35%

The feature pipeline and the training code computed the same features two subtly different ways, and the gap quietly cost us accuracy nobody could explain. I unified them behind one definition, so what the model learned and what it saw in production were finally the same thing.

Reworked features and validation; offline accuracy rose from 13% to 33% and held online.

SQLVector DatabasesRAG
2014

Cut feature/serving skew to zero

Labeling was slow, inconsistent, and the bottleneck on every model we wanted to ship. I built the tooling and the guidelines that made it fast and repeatable, and put the annotator disagreements to work as a signal instead of noise.

Unified training and serving features behind one definition; the accuracy gap nobody could explain simply disappeared.

PythonNLPSpark

Bring craft and measurement to work that usually gets neither.

Off the clock

Locked
Machine Learning
Locked
RAGGenerative AIPython
Locked
BigQuerySQL
Locked
PyTorch
Locked
Machine LearningPythonLangChain

In public

Python
Locked
PythonSnowflake
Locked

Practical bits

Timezones
Europe/Africa
Location
Poland
Languages
Czech, English
Domains
E-commerce
Employment
Full-time
Annual salary
Available with a company account
Social links
Available after unlocking

Reach out

Unlock with a company account to message directly and reveal:

  • Full name
  • Projects
  • In public
  • Social links

Find someone like this

Jump back to the board with this background already filtered in.