A novel improved random forest for text classification using feature ranking and optimal number of trees

Nasir Jalal; Arif Mehmood; Gyu Sang Choi; Imran Ashraf

Journal of King Saud University: Computer and Information Sciences (Jun 2022)

A novel improved random forest for text classification using feature ranking and optimal number of trees

Nasir Jalal,
Arif Mehmood,
Gyu Sang Choi,
Imran Ashraf

Affiliations

Nasir Jalal: Department of Computer Science and Information Technology, The Islamia University of Bahawalpur, Bahawalpur 63100, Pakistan
Arif Mehmood: Department of Computer Science and Information Technology, The Islamia University of Bahawalpur, Bahawalpur 63100, Pakistan
Gyu Sang Choi: Information and Communication Engineering, Yeungnam University, Gyeongsan 38541, Republic of Korea; Corresponding authors.
Imran Ashraf: Information and Communication Engineering, Yeungnam University, Gyeongsan 38541, Republic of Korea; Corresponding authors.

Journal volume & issue: Vol. 34, no. 6
pp. 2733 – 2742

Abstract

Read online

Machine learning-based models like random forest (RF) have been widely deployed in diverse domains such as image processing, health care, and text processing, etc. during the past few years. The RF is a prominent technique for handling imbalanced data and performs significantly better than other machine learning models due to its parallel architecture. This study presents an improved random forest for text classification, called improved random forest for text classification (IRFTC), that incorporates bootstrapping and random subspace methods simultaneously. The IRFTC removes unimportant (less important) features, adds a number of trees in the forest on each iteration, and monitors the classification performance of RF. Classification accuracy is determined with respect to the number of trees which defines the optimal number of trees for IRFTC. Feature ranking is determined using the quality of the split in a tree. The proposed IRFTC is applied on four different benchmark datasets, binary and multiclass, to validate its performance in this study. Results indicate that IRFTC outperforms both the traditional RF, as well as, other machine learning models such as logistic regression, support vector machine, Naive Bayes, and decision trees.

Published in Journal of King Saud University: Computer and Information Sciences

ISSN: 1319-1578 (Print)
Publisher: Elsevier
Country of publisher: Saudi Arabia
LCC subjects: Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: http://www.journals.elsevier.com/journal-of-king-saud-university-computer-and-information-sciences/

About the journal

Abstract

Keywords