Automatically Detect Software Security Vulnerabilities Based on Natural Language Processing Techniques and Machine Learning Algorithms

Do Xuan Cho; Vu Ngoc Son; Duong Duc

doi:10.5614/itbj.ict.res.appl.2022.16.1.5

Journal of ICT Research and Applications (May 2022)

Automatically Detect Software Security Vulnerabilities Based on Natural Language Processing Techniques and Machine Learning Algorithms

Do Xuan Cho,
Vu Ngoc Son,
Duong Duc

Affiliations

Do Xuan Cho: Faculty of Information Assurance, Posts and Telecommunications Institute of Technology, Hanoi, Vietnam
Vu Ngoc Son: Information Assurance Departement, FPT University, Hanoi, Vietnam
Duong Duc: Information Assurance Departement, FPT University, Hanoi, Vietnam

DOI: https://doi.org/10.5614/itbj.ict.res.appl.2022.16.1.5
Journal volume & issue: Vol. 16, no. 1

Abstract

Read online

Nowadays, software vulnerabilities pose a serious problem, because cyber-attackers often find ways to attack a system by exploiting software vulnerabilities. Detecting software vulnerabilities can be done using two main methods: i) signature-based detection, i.e. methods based on a list of known security vulnerabilities as a basis for contrasting and comparing; ii) behavior analysis-based detection using classification algorithms, i.e., methods based on analyzing the software code. In order to improve the ability to accurately detect software security vulnerabilities, this study proposes a new approach based on a technique of analyzing and standardizing software code and the random forest (RF) classification algorithm. The novelty and advantages of our proposed method are that to determine abnormal behavior of functions in the software, instead of trying to define behaviors of functions, this study uses the Word2vec natural language processing model to normalize and extract features of functions. Finally, to detect security vulnerabilities in the functions, this study proposes to use a popular and effective supervised machine learning algorithm.

Published in Journal of ICT Research and Applications

ISSN: 2337-5787 (Print); 2338-5499 (Online)
Publisher: ITB Journal Publisher
Country of publisher: Indonesia
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering: Telecommunication; Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: http://journals.itb.ac.id/index.php/jictra/index

About the journal

Abstract

Keywords