Sample profile — not a real person

Reworked features and validation; offline accuracy rose from 41% to 59% and held online.

See the full story

Proof of work

2022

Raised model accuracy 41%→56%

The feature pipeline and the training code computed the same features two subtly different ways, and the gap quietly cost us accuracy nobody could explain. I unified them behind one definition, so what the model learned and what it saw in production were finally the same thing.

Reworked features and validation; offline accuracy rose from 41% to 59% and held online.

SQLNLPscikit-learnSpark
2020

Added prediction monitoring across 5 models

Pipelines broke silently and nobody noticed for days, so the first sign of trouble was usually someone downstream asking why a report looked wrong. I rebuilt them with tests, explicit contracts between producers and consumers, and alerting that fired on the data rather than the job.

Instrumented live predictions for drift and degradation on 7 models; problems surfaced on a dashboard, not in support.

RAGSQL
2026

Made experiment analysis self-serve for 11 teams

We were flying blind on what the model did in production, with no idea it was wrong until a user complained loudly enough. I added monitoring on the predictions themselves, so drift and degradation showed up on a dashboard instead of in the support queue.

Built an A/B analysis layer; 10 product teams now read results without a data scientist in the loop.

dbtSQLPythonRAG
2019

Cut model inference cost 15% with quantization

Everyone wanted 'AI in the product' and nobody had said what good would look like, so we were one demo away from shipping something confidently wrong. I set the evaluation first — the metric, the holdout, the bar to clear — so we could tell real progress from a nice-looking demo.

Quantized and batched serving; cost per 1k requests fell 45% at the same quality bar.

Deep LearningPythonSQLMachine Learning
2018

Shipped a RAG assistant with 30% fewer hallucinations

Search results were mediocre and clickthrough had been flat for a year, with nobody owning the number end to end. I took relevance from the ranking model through the evaluation set to the online test, and made every change defensible against a holdout instead of a hunch.

Grounded answers in retrieval with citation checks; hallucination rate dropped 50% against the eval set.

TensorFlowscikit-learnPyTorchVector Databases

Bridge product intent and engineering reality without losing either.

Off the clock

Locked
PythonLLMsSpark
Locked
SQLNLPPyTorch
Locked
AirflowSQLGenerative AI
Locked
SparkMachine LearningPython
Locked
PythonSQLSpark

In public

Generative AIPythonLangChain
Locked
SQLSpark
Locked

Practical bits

Timezones
Europe/Africa
Location
Poland
Languages
Portuguese, Polish
Domains
Government, Social Media
Employment
Part-time · Full-time
Annual salary
Available with a company account
Social links
Available after unlocking

Reach out

Unlock with a company account to message directly and reveal:

  • Full name
  • Projects
  • In public
  • Social links

Find someone like this

Jump back to the board with this background already filtered in.