Cooperative Hybrid Semi-Supervised Learning for Text Sentiment Classification

Yang Li; Ying Lv; Suge Wang; Jiye Liang; Juanzi Li; Xiaoli Li

doi:10.3390/sym11020133

Symmetry (Jan 2019)

Cooperative Hybrid Semi-Supervised Learning for Text Sentiment Classification

Yang Li,
Ying Lv,
Suge Wang,
Jiye Liang,
Juanzi Li,
Xiaoli Li

Affiliations

Yang Li: School of Computer and Information Technology, Shanxi University, Taiyuan 030006, China
Ying Lv: School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China
Suge Wang: School of Computer and Information Technology, Shanxi University, Taiyuan 030006, China
Jiye Liang: School of Computer and Information Technology, Shanxi University, Taiyuan 030006, China
Juanzi Li: Computer Science Department, Tsinghua University, Beijing 100084, China
Xiaoli Li: Institute for Infocomm Research, A*Star, Singapore 138632, Singapore

DOI: https://doi.org/10.3390/sym11020133
Journal volume & issue: Vol. 11, no. 2
p. 133

Abstract

Read online

A large-scale and high-quality training dataset is an important guarantee to learn an ideal classifier for text sentiment classification. However, manually constructing such a training dataset with sentiment labels is a labor-intensive and time-consuming task. Therefore, based on the idea of effectively utilizing unlabeled samples, a synthetical framework that covers the whole process of semi-supervised learning from seed selection, iterative modification of the training text set, to the co-training strategy of the classifier is proposed in this paper for text sentiment classification. To provide an important basis for selecting the seed texts and modifying the training text set, three kinds of measures—the cluster similarity degree of an unlabeled text, the cluster uncertainty degree of a pseudo-label text to a learner, and the reliability degree of a pseudo-label text to a learner—are defined. With these measures, a seed selection method based on Random Swap clustering, a hybrid modification method of the training text set based on active learning and self-learning, and an alternately co-training strategy of the ensemble classifier of the Maximum Entropy and Support Vector Machine are proposed and combined into our framework. The experimental results on three Chinese datasets (COAE2014, COAE2015, and a Hotel review, respectively) and five English datasets (Books, DVD, Electronics, Kitchen, and MR, respectively) in the real world verify the effectiveness of the proposed framework.

Published in Symmetry

ISSN: 2073-8994 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science: Mathematics
Website: http://www.mdpi.com/journal/symmetry/

About the journal

Abstract

Keywords