Discriminating Emotions in the Valence Dimension from Speech Using Timbre Features

Anvarjon Tursunov; Soonil Kwon; Hee-Suk Pang

doi:10.3390/app9122470

Applied Sciences (Jun 2019)

Discriminating Emotions in the Valence Dimension from Speech Using Timbre Features

Anvarjon Tursunov,
Soonil Kwon,
Hee-Suk Pang

Affiliations

Anvarjon Tursunov: Department of Digital Contents, Sejong University, Seoul 05006, Korea
Soonil Kwon: Department of Digital Contents, Sejong University, Seoul 05006, Korea
Hee-Suk Pang: Department of Electrical Engineering, Sejong University, Seoul 05006, Korea

DOI: https://doi.org/10.3390/app9122470
Journal volume & issue: Vol. 9, no. 12
p. 2470

Abstract

Read online

The most used and well-known acoustic features of a speech signal, the Mel frequency cepstral coefficients (MFCC), cannot characterize emotions in speech sufficiently when a classification is performed to classify both discrete emotions (i.e., anger, happiness, sadness, and neutral) and emotions in valence dimension (positive and negative). The main reason for this is that some of the discrete emotions, such as anger and happiness, share similar acoustic features in the arousal dimension (high and low) but are different in the valence dimension. Timbre is a sound quality that can discriminate between two sounds even with the same pitch and loudness. In this paper, we analyzed timbre acoustic features to improve the classification performance of discrete emotions as well as emotions in the valence dimension. Sequential forward selection (SFS) was used to find the most relevant acoustic features among timbre acoustic features. The experiments were carried out on the Berlin Emotional Speech Database and the Interactive Emotional Dyadic Motion Capture Database. Support vector machine (SVM) and long short-term memory recurrent neural network (LSTM-RNN) were used to classify emotions. The significant classification performance improvements were achieved using a combination of baseline and the most relevant timbre acoustic features, which were found by applying SFS on a classification of emotions for the Berlin Emotional Speech Database. From extensive experiments, it was found that timbre acoustic features could characterize emotions sufficiently in a speech in the valence dimension.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords