I am a machine learning researcher based in Madrid. I work on retrieval, probabilistic models and learning on graphs, and I have spent the last eleven years building models that run in production.
Most of my work sits between two things that are usually kept apart: methods that are worth publishing, and systems that have to run every day without breaking. The domains have changed over the years, from parliamentary archives to protein networks to credit and fraud, but the underlying problems have not moved much. They are nearly always about ranking, similarity, and learning from partial labels.
This site holds my publications, the software I have released, and a growing set of technical notes.
Background
I read Computer Science at the University of Granada and stayed there for my doctorate, completed in 2010 and funded by an FPU grant from the Spanish Ministry of Science, awarded that year to thirty five students across the whole of computer science, electronics and communications. My thesis, supervised by L. M. de Campos and J. M. Fernández-Luna, was on document classification models based on Bayesian networks: automatic indexing against a thesaurus, classification in linked environments, and retrieval over structured documents. In 2008 I spent three months at the Laboratoire d'Informatique de Paris 6 with Patrick Gallinari and Ludovic Denoyer.
From 2010 to 2015 I was a postdoctoral researcher at Royal Holloway, University of London, in Alberto Paccanaro's lab, working on graph-based methods for protein function prediction, on consensus domain architecture for annotating protein families, and on semantic similarity between phenotypes. In 2013 I spent two months as a visiting fellow at Cornell University, hosted by Haiyuan Yu, where I developed the clustering method behind mutation3D, a tool that predicts cancer genes from the spatial arrangement of coding variants on protein structures. I am a co-first author of the resulting paper.
In industry
I joined Experian's UK&I DataLab in London in 2015 as its third scientist, working on business growth models, late payment prediction and graph-based application fraud detection. From 2017 to 2019 I was at BBVA in Madrid, the first Python and Spark specialist brought into my department to move it off SAS and onto the bank's private cloud, and I mentored several colleagues through that transition.
I returned to Experian in 2019 as a Senior Data Scientist in the UK&I and EMEA DataLab, working on privacy-preserving learning, algorithmic fairness, credit scoring with alternative data and fraud models for transactional data. Since late 2024 I have been an Innovation Manager, still hands on. I am also reading for a BSc in Mathematics at the UNED, part time and at a deliberately slow pace.