Big Data Analytics for the ATLAS EventIndex Project with Apache Spark

Álvaro Fernández Casaní; Carlos García Montoro; Santiago González de la Hoz; José Salt; Javier Sánchez; Miguel Villaplana Pérez

doi:10.1155/2023/6900908

Computational and Mathematical Methods (Jan 2023)

Big Data Analytics for the ATLAS EventIndex Project with Apache Spark

Álvaro Fernández Casaní,
Carlos García Montoro,
Santiago González de la Hoz,
José Salt,
Javier Sánchez,
Miguel Villaplana Pérez

Affiliations

Álvaro Fernández Casaní: Institute of Corpuscular Physics-IFIC (CSIC/UV)
Carlos García Montoro: Institute of Corpuscular Physics-IFIC (CSIC/UV)
Santiago González de la Hoz: Institute of Corpuscular Physics-IFIC (CSIC/UV)
José Salt: Institute of Corpuscular Physics-IFIC (CSIC/UV)
Javier Sánchez: Institute of Corpuscular Physics-IFIC (CSIC/UV)
Miguel Villaplana Pérez: Institute of Corpuscular Physics-IFIC (CSIC/UV)

DOI: https://doi.org/10.1155/2023/6900908
Journal volume & issue: Vol. 2023

Abstract

Read online

The ATLAS EventIndex was designed to provide a global event catalogue and limited event-level metadata for ATLAS experiment of the Large Hadron Collider (LHC) and their analysis groups and users during Run 2 (2015-2018) and has been running in production since. The LHC Run 3, started in 2022, has seen increased data-taking and simulation production rates, with which the current infrastructure would still cope but may be stretched to its limits by the end of Run 3. A new core storage service is being developed in HBase/Phoenix, and there is work in progress to provide at least the same functionality as the current one for increased data ingestion and search rates and with increasing volumes of stored data. In addition, new tools are being developed for solving the needed access cases within the new storage. This paper describes a new tool using Spark and implemented in Scala for accessing the big data quantities of the EventIndex project stored in HBase/Phoenix. With this tool, we can offer data discovery capabilities at different granularities, providing Spark Dataframes that can be used or refined within the same framework. Data analytic cases of the EventIndex project are implemented, like the search for duplicates of events from the same or different datasets. An algorithm and implementation for the calculation of overlap matrices of events across different datasets are presented. Our approach can be used by other higher-level tools and users, to ease access to the data in a performant and standard way using Spark abstractions. The provided tools decouple data access from the actual data schema, which makes it convenient to hide complexity and possible changes on the backed storage.

Published in Computational and Mathematical Methods

ISSN: 2577-7408 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Science: Mathematics; Technology: Technology (General): Industrial engineering. Management engineering: Applied mathematics. Quantitative methods
Website: https://onlinelibrary.wiley.com/journal/cmm

About the journal