Revisiting K-Means and Topic Modeling, a Comparison Study to Cluster Arabic Documents

M. Alhawarat; M. Hegazi

doi:10.1109/ACCESS.2018.2852648

IEEE Access (Jan 2018)

Revisiting K-Means and Topic Modeling, a Comparison Study to Cluster Arabic Documents

M. Alhawarat,
M. Hegazi

Affiliations

M. Alhawarat: ORCiD; Department of Computer Science, Prince Sattam Bin Abdulaziz University, Al-Kharj, Saudi Arabia
M. Hegazi: ORCiD; Department of Computer Science, Prince Sattam Bin Abdulaziz University, Al-Kharj, Saudi Arabia

DOI: https://doi.org/10.1109/ACCESS.2018.2852648
Journal volume & issue: Vol. 6
pp. 42740 – 42749

Abstract

Read online

Clustering Arabic text documents is of high importance for many natural language technologies. This paper uses a combined method to cluster Arabic text documents. Mainly, we use generative models and clustering techniques. The study uses latent Dirichlet allocation and k-means clustering algorithm and applies them to a news data set used in previous similar studies. The aim of this paper is twofold: it first shows that normalizing the weights in the vector space, for the document-term matrix of the text documents, dramatically improves the quality of clusters and hence the accuracy of clustering when using k-means algorithm. The results are compared to a recent study on clustering Arabic text documents. Second, it shows that the combined method is superior in terms of clustering quality for Arabic text documents according to external measures, such as purity, F-measure, entropy, accuracy, and other measures. It is shown in this paper that the purity of the combined method is 0.933 compared to 0.82 for k-means algorithm, and these figures are higher in comparison to a recent similar study. This is also confirmed by the other used validation measures. The correctness of the combined method is then confirmed using different Arabic data sets.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords