about
Pipes: A Meta-Dataset of Machine Learning Pipelines (arxiv.org)
3 points by gidellav on Sep 14, 2025 | hide | past | pdf | discuss on HN

In plain words: PIPES is a collection of results from running every combination of data-preparation and learning steps on 300 datasets, so tools can pick the best pipeline for a new dataset cheaply. Unlike OpenML's records, which favor a few popular steps, it covers 9,408 pipelines evenly.

Abstract · PIPES: A Meta-dataset of Machine Learning Pipelines

Solutions to the Algorithm Selection Problem (ASP) in machine learning face the challenge of high computational costs associated with evaluating various algorithms' performances on a given dataset. To mitigate this cost, the meta-learning field can leverage previously executed experiments shared in online repositories such as OpenML. OpenML provides an extensive collection of machine learning experiments. However, an analysis of OpenML's records reveals limitations. It lacks diversity in pipelines, specifically when exploring data preprocessing steps/blocks, such as scaling or imputation, resulting in limited representation. Its experiments are often focused on a few popular techniques within each pipeline block, leading to an imbalanced sample. To overcome the observed limitations of OpenML, we propose PIPES, a collection of experiments involving multiple pipelines designed to represent all combinations of the selected sets of techniques, aiming at diversity and completeness. PIPES stores the results of experiments performed applying 9,408 pipelines to 300 datasets. It includes detailed information on the pipeline blocks, training and testing times, predictions, performances, and the eventual error messages. This comprehensive collection of results allows researchers to perform analyses across diverse and representative pipelines and datasets. PIPES also offers potential for expansion, as additional data and experiments can be incorporated to support the meta-learning community further. The data, code, supplementary material, and all experiments can be found at https://github.com/cynthiamaia/PIPES.git.

Cynthia Moreira Maia, Lucas B. V. de Amorim, George D. C. Cavalcanti, Rafael M. O. Cruz
arXiv:2509.09512 · cs.LG · submitted Sep 11, 2025
abstract · pdf · html

add comment on HN