Bimodal Emotion Recognition Model for Minnan Songs

Zhenglong Xiang; Xialei Dong; Yuanxiang Li; Fei Yu; Xing Xu; Hongrun Wu

doi:10.3390/info11030145

Information (Mar 2020)

Bimodal Emotion Recognition Model for Minnan Songs

Zhenglong Xiang,
Xialei Dong,
Yuanxiang Li,
Fei Yu,
Xing Xu,
Hongrun Wu

Affiliations

Zhenglong Xiang: School of Computer Science, Wuhan University, Wuhan 430072, China
Xialei Dong: School of Computer Science, Wuhan University, Wuhan 430072, China
Yuanxiang Li: School of Computer Science, Wuhan University, Wuhan 430072, China
Fei Yu: School of Physics and Information Engineering, Minnan Normal University, Zhangzhou 363000, China
Xing Xu: School of Physics and Information Engineering, Minnan Normal University, Zhangzhou 363000, China
Hongrun Wu: School of Physics and Information Engineering, Minnan Normal University, Zhangzhou 363000, China

DOI: https://doi.org/10.3390/info11030145
Journal volume & issue: Vol. 11, no. 3
p. 145

Abstract

Read online

Most of the existing research papers study the emotion recognition of Minnan songs from the perspectives of music analysis theory and music appreciation. However, these investigations do not explore any possibility of carrying out an automatic emotion recognition of Minnan songs. In this paper, we propose a model that consists of four main modules to classify the emotion of Minnan songs by using the bimodal data—song lyrics and audio. In the proposed model, an attention-based Long Short-Term Memory (LSTM) neural network is applied to extract lyrical features, and a Convolutional Neural Network (CNN) is used to extract the audio features from the spectrum. Then, two kinds of extracted features are concatenated by multimodal compact bilinear pooling, and finally, the concatenated features are input to the classifying module to determine the song emotion. We designed three experiment groups to investigate the classifying performance of combinations of the four main parts, the comparisons of proposed model with the current approaches and the influence of a few key parameters on the performance of emotion recognition. The results show that the proposed model exhibits better performance over all other experimental groups. The accuracy, precision and recall of the proposed model exceed 0.80 in a combination of appropriate parameters.

Published in Information

ISSN: 2078-2489 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: http://www.mdpi.com/journal/information/

About the journal

Abstract

Keywords