Data engineering · Multilingual NLP
Align and Shine
A traceable pipeline that transforms noisy multilingual source data into aligned corpora across Catalan, English, French, Italian and Spanish.
Cut a core workflow from 48 to 6 hours and released the resulting corpus as an open, reproducible data product.