Smart Data Placement Using Storage-as-a-Service Model for Big Data Pipelines

Akif Quddus Khan; Nikolay Nikolov; Mihhail Matskin; Radu Prodan; Dumitru Roman; Bekir Sahin; Christoph Bussler; Ahmet Soylu

doi:10.3390/s23020564

Sensors (Jan 2023)

Smart Data Placement Using Storage-as-a-Service Model for Big Data Pipelines

Akif Quddus Khan,
Nikolay Nikolov,
Mihhail Matskin,
Radu Prodan,
Dumitru Roman,
Bekir Sahin,
Christoph Bussler,
Ahmet Soylu

Affiliations

Akif Quddus Khan: Department of Computer Science, Norwegian University of Science and Technology—NTNU, 2815 Gjøvik, Norway
Nikolay Nikolov: SINTEF Digital, SINTEF AS, 0373 Oslo, Norway
Mihhail Matskin: Department of Computer Science, KTH Royal Institute of Technology, 114 28 Stockholm, Sweden
Radu Prodan: Department of Information Technology, University of Klagenfurt, 9020 Klagenfurt, Austria
Dumitru Roman: SINTEF Digital, SINTEF AS, 0373 Oslo, Norway
Bekir Sahin: Logistics Management, National University of Science and Technology, 111 Sohar, Oman
Christoph Bussler: Robert Bosch LLC, Sunnyvale, CA 94085, USA
Ahmet Soylu: Department of Computer Science, OsloMet—Oslo Metropolitan University, 0167 Oslo, Norway

DOI: https://doi.org/10.3390/s23020564
Journal volume & issue: Vol. 23, no. 2
p. 564

Abstract

Read online

Big data pipelines are developed to process data characterized by one or more of the three big data features, commonly known as the three Vs (volume, velocity, and variety), through a series of steps (e.g., extract, transform, and move), making the ground work for the use of advanced analytics and ML/AI techniques. Computing continuum (i.e., cloud/fog/edge) allows access to virtually infinite amount of resources, where data pipelines could be executed at scale; however, the implementation of data pipelines on the continuum is a complex task that needs to take computing resources, data transmission channels, triggers, data transfer methods, integration of message queues, etc., into account. The task becomes even more challenging when data storage is considered as part of the data pipelines. Local storage is expensive, hard to maintain, and comes with several challenges (e.g., data availability, data security, and backup). The use of cloud storage, i.e., storage-as-a-service (StaaS), instead of local storage has the potential of providing more flexibility in terms of scalability, fault tolerance, and availability. In this article, we propose a generic approach to integrate StaaS with data pipelines, i.e., computation on an on-premise server or on a specific cloud, but integration with StaaS, and develop a ranking method for available storage options based on five key parameters: cost, proximity, network performance, server-side encryption, and user weights/preferences. The evaluation carried out demonstrates the effectiveness of the proposed approach in terms of data transfer performance, utility of the individual parameters, and feasibility of dynamic selection of a storage option based on four primary user scenarios.

Published in Sensors

ISSN: 1424-8220 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Chemical technology
Website: http://www.mdpi.com/journal/sensors

About the journal

Abstract

Keywords