Ioannis Koumarelas, PhD

Ioannis Koumarelas, PhD

Machine Learning Engineer
PhD in Data Quality

Mission Statement

Data scientist (PhD) with 5+ years of experience building and deploying production ML systems. Experience designing scalable pipelines that process billions of data points and turning research prototypes into production-grade systems. Deep expertise in data quality, entity resolution, and duplicate detection. Production experience with embeddings and vector search, plus recent work in LLM application development and agentic systems.

Experience

Senior Data Scientist / Data Scientist

Veeva Systems – Link Product

Senior Data Scientist (Mar 2024 – Feb 2026) · Data Scientist (Dec 2021 – Feb 2024)

  • Was part of a team that built the ML models matching and clustering medical activities into expert profiles – embedding-based candidate retrieval, similarity ranking with gradient boosted trees, and entity resolution – processing up to 20 billion activity pairs per run and generating millions of automated profiles across US, EU, LATAM, and APAC regions.
  • Transformed exploratory Jupyter Notebook prototypes into production-ready PySpark + Airflow pipelines on AWS EMR, with MLflow for experiment tracking and model deployment, Docker and Kubernetes for containerized services, testing, monitoring, and CI/CD integration, collaborating with cross-functional engineering teams.
  • Balanced a three-way trade-off between precision, recall, and manual-curation cost, using threshold-based quality tiers to hold 99% precision while reducing manual curation costs by up to 70%.
  • Organized Data Science meetups, technical talks, and team activities to promote knowledge sharing and strengthen engineering culture.

Data Engineer / Full-Stack Engineer

HPI Schul-Cloud – Dataport
  • Built and maintained data pipelines for 300k+ educational assets, improving structure, reliability, and discoverability for end users.
  • Implemented systematic data preparation, cleaning workflows, and duplicate-detection methods to ensure data quality at scale.
  • Contributed across the full stack (Python, Vue.js, PostgreSQL, Docker, Kubernetes) to maintain and scale the educational platform.
  • Led technical requirements clarification, team operations, and onboarding during a multi-month organizational transition.

Research Consultant

SAP & SAP Concur
  • Developed 3 novel ML pipelines in Python and Java to improve duplicate detection, increasing matching success by 18%.
  • Delivered on-site technical tutorials at SAP Concur Seattle (USA) on data matching classification and pipeline optimization.

Technical Skills

Programming Languages

Python SQL Java JavaScript C/C++

ML & Data

PySpark scikit-learn XGBoost PyTorch Pandas NumPy Hugging Face Transformers MLflow Apache Spark Apache Airflow

Infrastructure & DevOps

AWS (EMR, S3) Docker Kubernetes CI/CD FastAPI Git pytest

Databases

PostgreSQL MongoDB

Data Quality

Entity Resolution Duplicate Detection Record Linkage Data Cleaning Data Preparation Embeddings Vector Search (Qdrant)

AI & LLMs

Large Language Models LangChain LangGraph Agentic AI

Certificates

AI & LLM Engineering (Udemy)
Udemy ∙ October 2025

Completed during July - October 2025 a comprehensive series of courses covering modern AI engineering practices:

Courses completed:

Generative AI with Large Language Models
Coursera ∙ July 2025
Three-week course covering the complete LLM lifecycle: Transformer architecture and pretraining, fine-tuning techniques including Parameter Efficient Fine-Tuning (PEFT) with LoRA and Soft Prompts, and Reinforcement Learning with Human Feedback (RLHF). Explored Chain-of-Thought reasoning and the ReAct framework that underlies modern agentic AI systems.
Deep Learning Specialization
Coursera ∙ January 2021

Foundational specialization from Coursera on Deep Learning. Comprised of the following courses:

  1. Neural Networks and Deep Learning
  2. Improving Deep Neural Networks: Hyperparameter Tuning, Regularization and Optimization
  3. Structuring Machine Learning Projects
  4. Convolutional Neural Networks
  5. Sequence Models

Through it I got a hollistic refreshment and further expansion of my knowledge on the primary Deep Learning fundamentals and models.

Education

Intensive German Course – Levels A2.2, B1.1, B1.2

Die Neue Schule, Berlin
Intensive German language course in Berlin, progressing through levels A2.2, B1.1, and B1.2.

PhD (Dr. rer. nat.) in Information Systems – Data Preparation & Domain-Agnostic Duplicate Detection

Hasso Plattner Institute (HPI), University of Potsdam – Digital Engineering Faculty
Thesis on Data Preparation and Domain-Agnostic Duplicate Detection, supervised by Prof. Felix Naumann. Defended with distinction (Magna cum Laude). Published 7 papers in top-tier journals and conferences. Organized 6 project seminars on Duplicate Detection, Data Preparation, Blockchain, Text Mining, and Recommender Systems.
Read dissertation

MSc in Informatics (Information Systems) – Theta-Joins on MapReduce

Aristotle University of Thessaloniki (AUTH), School of Informatics
Implemented thesis in Python, Java, and Hadoop; published in top-tier conference. Awarded State Scholarship Foundation scholarship. Vice Chair of local ACM Student Chapter. Participated in ACM SIGMOD 2013 programming contest (streaming system in C++).
Read thesis

BSc in Informatics (Information Systems) – Recommender System on MapReduce

Aristotle University of Thessaloniki (AUTH), School of Informatics
Implemented thesis in Java and Hadoop; published in top-tier journal. Interned at IT Center performing system and database administration.
Read thesis (in Greek)
Publications

10 peer-reviewed papers across BSc, MSc, and PhD research — 6 journal articles and 4 conference papers on entity resolution, duplicate detection, data preparation, and data quality.

Journals — ACM Journal of Data and Information Quality (×3) · ACM Transactions on Database Systems · Distributed and Parallel Databases (Springer) · Expert Systems with Applications (Elsevier)

Conferences — VLDB · EDBT/ICDT · Italian Symposium on Advanced Database Systems (SEBD) · International Workshop on Data Science for Macro-Modeling

Browse all publications · Google Scholar profile

Languages

🇬🇷 Greek Native
🇬🇧 English Fluent
🇩🇪 German Intermediate (B1)