Quality assurance strategies for machine learning applications in big data analytics: an overview

Mihajlo Ogrizović; Dražen Drašković; Dragan Bojić

doi:10.1186/s40537-024-01028-y

Journal of Big Data (Oct 2024)

Quality assurance strategies for machine learning applications in big data analytics: an overview

Mihajlo Ogrizović,
Dražen Drašković,
Dragan Bojić

Affiliations

Mihajlo Ogrizović: School of Electrical Engineering, Department of Computer Science and Information Technology, University of Belgrade
Dražen Drašković: School of Electrical Engineering, Department of Computer Science and Information Technology, University of Belgrade
Dragan Bojić: School of Electrical Engineering, Department of Computer Science and Information Technology, University of Belgrade

DOI: https://doi.org/10.1186/s40537-024-01028-y
Journal volume & issue: Vol. 11, no. 1
pp. 1 – 48

Abstract

Read online

Abstract Machine learning (ML) models have gained significant attention in a variety of applications, from computer vision to natural language processing, and are almost always based on big data. There are a growing number of applications and products with built-in machine learning models, and this is the area where software engineering, artificial intelligence and data science meet. The requirement for a system to operate in a real-world environment poses many challenges, such as how to design for wrong predictions the model may make; How to assure safety and security despite possible mistakes; which qualities matter beyond a model’s prediction accuracy; How can we identify and measure important quality requirements, including learning and inference latency, scalability, explainability, fairness, privacy, robustness, and safety. It has become crucial to test thoroughly these models to assess their capabilities and potential errors. Existing software testing methods have been adapted and refined to discover faults in machine learning and deep learning models. This paper covers a taxonomy, a methodologically uniform presentation of all presented solutions to the aforementioned issues, as well as conclusions about possible future development trends. The main contributions of this paper are a classification that closely follows the structure of the ML-pipeline, a precisely defined role of each team member within that pipeline, an overview of trends and challenges in the combination of ML and big data analytics, with uses in the domains of industry and education.

Published in Journal of Big Data

ISSN: 2196-1115 (Online)
Publisher: SpringerOpen
Country of publisher: United Kingdom
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering: Electronics: Computer engineering. Computer hardware; Technology: Technology (General): Industrial engineering. Management engineering: Information technology; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://journalofbigdata.springeropen.com

About the journal

Abstract

Keywords