Block size estimation for data partitioning in HPC applications using machine learning techniques

Riccardo Cantini; Fabrizio Marozzo; Alessio Orsino; Domenico Talia; Paolo Trunfio; Rosa M. Badia; Jorge Ejarque; Fernando Vázquez-Novoa

doi:10.1186/s40537-023-00862-w

Journal of Big Data (Jan 2024)

Block size estimation for data partitioning in HPC applications using machine learning techniques

Riccardo Cantini,
Fabrizio Marozzo,
Alessio Orsino,
Domenico Talia,
Paolo Trunfio,
Rosa M. Badia,
Jorge Ejarque,
Fernando Vázquez-Novoa

Affiliations

Riccardo Cantini: University of Calabria
Fabrizio Marozzo: University of Calabria
Alessio Orsino: University of Calabria
Domenico Talia: University of Calabria
Paolo Trunfio: University of Calabria
Rosa M. Badia: Barcelona Supercomputing Center
Jorge Ejarque: Barcelona Supercomputing Center
Fernando Vázquez-Novoa: Barcelona Supercomputing Center

DOI: https://doi.org/10.1186/s40537-023-00862-w
Journal volume & issue: Vol. 11, no. 1
pp. 1 – 23

Abstract

Read online

Abstract The extensive use of HPC infrastructures and frameworks for running data-intensive applications has led to a growing interest in data partitioning techniques and strategies. In fact, application performance can be heavily affected by how data are partitioned, which in turn depends on the selected size for data blocks, i.e. the block size. Therefore, finding an effective partitioning, i.e. a suitable block size, is a key strategy to speed-up parallel data-intensive applications and increase scalability. This paper describes a methodology, namely BLEST-ML (BLock size ESTimation through Machine Learning), for block size estimation that relies on supervised machine learning techniques. The proposed methodology was evaluated by designing an implementation tailored to dislib, a distributed computing library highly focused on machine learning algorithms built on top of the PyCOMPSs framework. We assessed the effectiveness of the provided implementation through an extensive experimental evaluation considering different algorithms from dislib, datasets, and infrastructures, including the MareNostrum 4 supercomputer. The results we obtained show the ability of BLEST-ML to efficiently determine a suitable way to split a given dataset, thus providing a proof of its applicability to enable the efficient execution of data-parallel applications in high performance environments.

Published in Journal of Big Data

ISSN: 2196-1115 (Online)
Publisher: SpringerOpen
Country of publisher: United Kingdom
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering: Electronics: Computer engineering. Computer hardware; Technology: Technology (General): Industrial engineering. Management engineering: Information technology; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://journalofbigdata.springeropen.com

About the journal

Abstract

Keywords