MLMD—A Malware-Detecting Antivirus Tool Based on the XGBoost Machine Learning Algorithm

Jakub Palša; Norbert Ádám; Ján Hurtuk; Eva Chovancová; Branislav Madoš; Martin Chovanec; Stanislav Kocan

doi:10.3390/app12136672

Applied Sciences (Jul 2022)

MLMD—A Malware-Detecting Antivirus Tool Based on the XGBoost Machine Learning Algorithm

Jakub Palša,
Norbert Ádám,
Ján Hurtuk,
Eva Chovancová,
Branislav Madoš,
Martin Chovanec,
Stanislav Kocan

Affiliations

Jakub Palša: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia
Norbert Ádám: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia
Ján Hurtuk: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia
Eva Chovancová: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia
Branislav Madoš: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia
Martin Chovanec: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia
Stanislav Kocan: Department of Computers and Informatics, Faculty of Electrical Engineering and Informatics, Technical University of Košice, Letná 9, 042 00 Kosice, Slovakia

DOI: https://doi.org/10.3390/app12136672
Journal volume & issue: Vol. 12, no. 13
p. 6672

Abstract

Read online

This paper focuses on training machine learning models using the XGBoost and extremely randomized trees algorithms on two datasets obtained using static and dynamic analysis of real malicious and benign samples. We then compare their success rates—both mutually and with other algorithms, such as the random forest, the decision tree, the support vector machine, and the naïve Bayes algorithms, which we compared in our previous work on the same datasets. The best performing classification models, using the XGBoost algorithm, achieved 91.9% detection accuracy and 98.2% sensitivity, 0.853 AUC, and 0.949 F1 score on the static analysis dataset, and 96.4% accuracy and 98.5% sensitivity, 0.940 AUC, and 0.977 F1 score on the dynamic analysis dataset. Then, we exported the best performing machine learning models and used them in our proposed MLMD program, automating the process of static and dynamic analysis and allowing the trained models to be used for classification on new samples.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords