Data & AI Engineer · Leeds, UK

I build reliable data and AI systems for real-world decisions.

Data pipelines, machine-learning workflows and applied AI—designed around measurable outcomes, reproducibility and production constraints.

Based in Leeds, with 6+ years across banking, consulting, software and applied NLP. Open to Data Engineering, ML Engineering and applied AI roles.

6+years delivering data and software systems
50%lower processing latency in corporate banking
48 → 6hcore data-workflow runtime
DistinctionMSc Data Science & Analytics

Selected work

Evidence, not just a list of tools.

Each project starts with a concrete problem, makes the technical decisions visible and links to the underlying work.

Data engineering · Multilingual NLP

Align and Shine

A traceable pipeline that transforms noisy multilingual source data into aligned corpora across Catalan, English, French, Italian and Spanish.

Cut a core workflow from 48 to 6 hours and released the resulting corpus as an open, reproducible data product.

  • Python
  • Transformers
  • GPU / HPC
  • Data pipelines

Applied ML · Model risk

Late refill risk modelling

A leakage-aware temporal modelling pipeline for prescription refill risk, with calibration and explicit analysis of dataset shift.

Surfaced a material performance drop on the later test period—evidence that the model should not be deployed without addressing temporal shift.

  • Python
  • Scikit-learn
  • XGBoost
  • Temporal validation

Forecasting · Decision science

Exchange-rate forecasting

A simulation-based, out-of-sample comparison of statistical, structural and machine-learning approaches to foreign-exchange forecasting.

Found that random walks remained difficult to beat at short horizons, while structural hybrids improved longer-horizon forecasts for selected currency pairs.

  • Python
  • ARIMA
  • GARCH
  • Monte Carlo
View all projects

How I work

From raw data to reliable decisions.

I combine engineering discipline, statistical thinking and clear communication across the full lifecycle.

01

Data systems

Designing traceable, maintainable pipelines across batch processing, distributed compute and cloud workflows.

Python · SQL · PySpark · Hadoop · Airflow · AWS
02

Applied machine learning

Turning ambiguous problems into measurable experiments, careful validation and systems that support real decisions.

Scikit-learn · XGBoost · TensorFlow · MLflow
03

NLP & generative AI

Building multilingual corpora, retrieval workflows and language systems with reproducibility and evaluation in mind.

Transformers · RAG · LangGraph · Semantic search

Experience

Experience where data has to work.

Banking, consulting, software and applied AI—connected by a focus on dependable systems and useful decisions.

2024

Data Engineer

BBVA · Corporate & Investment Banking

Built PySpark and Hadoop pipelines orchestrated with AWS Step Functions, cutting processing latency by 50%, and integrated XGBoost classification through REST APIs.

2022 — 2024

Database Analyst

INDRA

Developed Python and SQL feature-engineering and ETL pipelines for late-payment and fraud modelling, translating model outputs into decision-ready KPIs.

2026 — present

Research Assistant · NLP & Data Science

University of Leeds

Build and optimise distributed NLP pipelines on enterprise Linux and NVIDIA GPU infrastructure, cutting a core multilingual workflow from 48 to 6 hours and releasing an open five-language data product.

Full experience

Beyond the code

A path through engineering, entrepreneurship and data.

From mechatronics and product-building in Peru to financial data systems and applied AI in Leeds.

Read my story