Big Data and Cognitive Computing (Mar 2022)

A Combined System Metrics Approach to Cloud Service Reliability Using Artificial Intelligence

  • Tek Raj Chhetri,
  • Chinmaya Kumar Dehury,
  • Artjom Lind,
  • Satish Narayana Srirama,
  • Anna Fensel

DOI
https://doi.org/10.3390/bdcc6010026
Journal volume & issue
Vol. 6, no. 1
p. 26

Abstract

Read online

Identifying and anticipating potential failures in the cloud is an effective method for increasing cloud reliability and proactive failure management. Many studies have been conducted to predict potential failure, but none have combined SMART (self-monitoring, analysis, and reporting technology) hard drive metrics with other system metrics, such as central processing unit (CPU) utilisation. Therefore, we propose a combined system metrics approach for failure prediction based on artificial intelligence to improve reliability. We tested over 100 cloud servers’ data and four artificial intelligence algorithms: random forest, gradient boosting, long short-term memory, and gated recurrent unit, and also performed correlation analysis. Our correlation analysis sheds light on the relationships that exist between system metrics and failure, and the experimental results demonstrate the advantages of combining system metrics, outperforming the state-of-the-art.

Keywords