Systematic Comparison of Vectorization Methods in Classification Context

Urszula Krzeszewska; Aneta Poniszewska-Marańda; Joanna Ochelska-Mierzejewska

doi:10.3390/app12105119

Applied Sciences (May 2022)

Systematic Comparison of Vectorization Methods in Classification Context

Urszula Krzeszewska,
Aneta Poniszewska-Marańda,
Joanna Ochelska-Mierzejewska

Affiliations

Urszula Krzeszewska: Institute of Information Technology, Lodz University of Technology, 93-590 Lodz, Poland
Aneta Poniszewska-Marańda: Institute of Information Technology, Lodz University of Technology, 93-590 Lodz, Poland
Joanna Ochelska-Mierzejewska: Institute of Information Technology, Lodz University of Technology, 93-590 Lodz, Poland

DOI: https://doi.org/10.3390/app12105119
Journal volume & issue: Vol. 12, no. 10
p. 5119

Abstract

Read online

Natural language processing has been the subject of numerous studies in the last decade. These have focused on the various stages of text processing, from text preparation to vectorization to final text comprehension. The goal of vector space modeling is to project words in a language corpus into a vector space in such a way that words that are similar in meaning are close to each other. Currently, there are two commonly used approaches to the topic of vectorization. The first focuses on creating word vectors taking into account the entire linguistic context, while the second focuses on creating document vectors in the context of the linguistic corpus of the analyzed texts. The paper presents the comparison of different existing text vectorization methods in natural language processing, especially in Text Mining. The comparison of text vectorization methods is possible by checking the accuracy of classification; we used the methods NBC and k-NN, as they are some of the simplest methods. They were used for the classification in order to avoid the influence of the choice of the method itself on the final result. The conducted experiments provide a basis for further research for better automatic text analysis.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords