COMPRESSED LEARNING FOR TEXT CATEGORIZATION

Artur Ferreira; Mario Figueiredo

ISEL Academic Journal of Electronics, Telecommunications and Computers (Jun 2013)

COMPRESSED LEARNING FOR TEXT CATEGORIZATION

Artur Ferreira,
Mario Figueiredo

Affiliations

Artur Ferreira
Mario Figueiredo

Journal volume & issue: Vol. 2, no. 1

Abstract

Read online

In text classification based on the bag-of-words (BoW) or similar representations, we usually have a large number of features, many of which are irrelevant (or even detrimental) for classification tasks. Recent results show that compressed learning (CL), i.e., learning in a domain of reduced dimensionality obtained by random projections (RP), is possible, and theoretical bounds on the test set error rate have been shown. In this work, we assess the performance of CL, based on RP of BoW representations for text classification. Our experimental results show that CL significantly reduces the number of features and the training time, while simultaneously improving the classification accuracy. Rather than the mild decrease in accuracy upper bounded by the theory, we actually find an increase of accuracy. Our approach is further compared against two techniques, namely the unsupervised random subspaces method and the supervised Fisher index. The CL approach is suited for unsupervised or semi-supervised learning, without any modification, since it does not use the class labels.

random projections, random subspaces, compressed learning, text classification, support vector machines

Published in ISEL Academic Journal of Electronics, Telecommunications and Computers

ISSN: 2182-4010 (Online)
Publisher: Instituto Superior de Engenharia de Lisboa (ISEL)
Country of publisher: Portugal
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering: Telecommunication; Technology: Electrical engineering. Electronics. Nuclear engineering: Electronics; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: http://journals.isel.pt/index.php/i-ETC/index

About the journal

Abstract

Keywords