CyBERT: Cybersecurity Claim Classification by Fine-Tuning the BERT Language Model

Kimia Ameri; Michael Hempel; Hamid Sharif; Juan Lopez Jr.; Kalyan Perumalla

doi:10.3390/jcp1040031

Journal of Cybersecurity and Privacy (Nov 2021)

CyBERT: Cybersecurity Claim Classification by Fine-Tuning the BERT Language Model

Kimia Ameri,
Michael Hempel,
Hamid Sharif,
Juan Lopez Jr.,
Kalyan Perumalla

Affiliations

Kimia Ameri: Department of Electrical & Computer Engineering, University of Nebraska-Lincoln, Lincoln, NE 68182, USA
Michael Hempel: Department of Electrical & Computer Engineering, University of Nebraska-Lincoln, Lincoln, NE 68182, USA
Hamid Sharif: Department of Electrical & Computer Engineering, University of Nebraska-Lincoln, Lincoln, NE 68182, USA
Juan Lopez Jr.: Oak Ridge National Laboratory, Oak Ridge, TN 37831, USA
Kalyan Perumalla: Oak Ridge National Laboratory, Oak Ridge, TN 37831, USA

DOI: https://doi.org/10.3390/jcp1040031
Journal volume & issue: Vol. 1, no. 4
pp. 615 – 637

Abstract

Read online

We introduce CyBERT, a cybersecurity feature claims classifier based on bidirectional encoder representations from transformers and a key component in our semi-automated cybersecurity vetting for industrial control systems (ICS). To train CyBERT, we created a corpus of labeled sequences from ICS device documentation collected across a wide range of vendors and devices. This corpus provides the foundation for fine-tuning BERT’s language model, including a prediction-guided relabeling process. We propose an approach to obtain optimal hyperparameters, including the learning rate, the number of dense layers, and their configuration, to increase the accuracy of our classifier. Fine-tuning all hyperparameters of the resulting model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original architecture to 94.4% obtained with CyBERT. Furthermore, we evaluated CyBERT for the impact of randomness in the initialization, training, and data-sampling phases. CyBERT demonstrated a standard deviation of ±0.6% during validation across 100 random seed values. Finally, we also compared the performance of CyBERT to other well-established language models including GPT2, ULMFiT, and ELMo, as well as neural network models such as CNN, LSTM, and BiLSTM. The results showed that CyBERT outperforms these models on the validation accuracy and the F1 score, validating CyBERT’s robustness and accuracy as a cybersecurity feature claims classifier.

Published in Journal of Cybersecurity and Privacy

ISSN: 2624-800X (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General)
Website: https://www.mdpi.com/journal/jcp

About the journal

Abstract

Keywords