An Investigation into the Process of Organizing and Retrieving Web Texts Based on the Integration of Semantic Concepts In order to organize knowledge

Saeede Anbaee Farimani; Hamid Tabatabaee; Mojtaba kaffashan kakhki

Iranian Journal of Information Processing & Management (Sep 2019)

An Investigation into the Process of Organizing and Retrieving Web Texts Based on the Integration of Semantic Concepts In order to organize knowledge

Saeede Anbaee Farimani,
Hamid Tabatabaee,
Mojtaba kaffashan kakhki

Affiliations

Saeede Anbaee Farimani: Mashhad Branch, Islamic Azad University, Mashhad, Iran
Hamid Tabatabaee: Quchan Branch, Islamic Azad University, Quchan, Iran
Mojtaba kaffashan kakhki: Ferdowsi University of Mashhad

Journal volume & issue: Vol. 34, no. 4
pp. 1879 – 1904

Abstract

Read online

Improvement in information retrieval performance relates to the method of knowledge extraction from large amounts of text information on web. Text classification is one of application of knowledge extraction with supervised machine learning methods. This paper proposed Kullback-Leibler divergence KNN for classifying extracted features based on term weighting with Latent Dirichlet Allocation Algorithm. LDA is Non Negative matrix factorization method proposed for topic modelling and dimension reduction of high dimensional feature space .In traditional LDA, each component value is assigned using the information retrieval TF measure, While this weighting method seems very appropriate for IR, it is not clear that it is the best choice for TC problems. Actually, this weighting method does not leverage the information implicitly contained in the categorization task to represent documents. In this paper, we introduce a new weighting method based on Point wise Mutual Information for accessing the importance of a word for a specific latent concept, then each document classified based on probability distribution over the latent topics. Experimental result investigated when we used PMI measure for term Weighing and KNN with Kullback-Leibler distance, accuracy has been 82.5%, with lower complexity and same accuracy versus complex deep learning methods.

Published in Iranian Journal of Information Processing & Management

ISSN: 2251-8223 (Print); 2251-8231 (Online)
Publisher: Iranian Research Institute for Information and Technology
Country of publisher: Iran, Islamic Republic of
LCC subjects: Bibliography. Library science. Information resources
Website: http://jipm.irandoc.ac.ir/index.php?slc_lang=en&sid=1

About the journal

Abstract

Keywords