SPARTA: Speaker Profiling for ARabic TAlk

Wael Farhan; Muhy Eddin Za'Ter; Qusai Abu Obaidah; Hisham Al Bataineh; Zyad Sober; Hussein Al Natsheh

doi:10.23919/FRUCT50888.2021.9347615

Proceedings of the XXth Conference of Open Innovations Association FRUCT (Jan 2021)

SPARTA: Speaker Profiling for ARabic TAlk

Wael Farhan,
Muhy Eddin Za'Ter,
Qusai Abu Obaidah,
Hisham Al Bataineh,
Zyad Sober,
Hussein Al Natsheh

Affiliations

Wael Farhan: Mawdoo3 Ltd, Jordan
Muhy Eddin Za'Ter: Mawdoo3 Ltd, Jordan
Qusai Abu Obaidah: Mawdoo3 Ltd, Jordan
Hisham Al Bataineh: Mawdoo3 Ltd, Jordan
Zyad Sober: Mawdoo3 Ltd, Jordan
Hussein Al Natsheh: Mawdoo3 Ltd, Jordan

DOI: https://doi.org/10.23919/FRUCT50888.2021.9347615
Journal volume & issue: Vol. 28, no. 1
pp. 103 – 110

Abstract

Read online

This paper proposes a novel approach to an automatic estimation of three speaker traits from Arabic speech: gender, emotion, and dialect. After showing promising results on different text classification tasks, the multi-task learning (MTL) approach is used in this paper for Arabic speech classification tasks. The dataset was assembled from six publicly available datasets. First, The datasets were edited and thoroughly divided into train, development, and test sets (open to the public), and a benchmark was set for each task and dataset throughout the paper. Then, three different networks were explored: Long Short Term Memory (LSTM), Convolutional Neural Network (CNN), and Fully-Connected Neural Network (FCNN) on five different types of features: two raw features (MFCC and MEL) and three pre-trained vectors (i-vectors, d-vectors, and x-vectors). LSTM and CNN networks were implemented using raw features: MFCC and MEL, wher FCNN was explored on the pre-trained vectors while varying the hyper-parameters of these networks to obtain the best results for each dataset and task. MTL was evaluated against the single task learning (STL) approach for the three tasks and six datasets, in which the MTL and pre-trained vectors almost constantly outperformed STL. All the data and pre-trained models used in this paper are available and can be acquired by the public.

Published in Proceedings of the XXth Conference of Open Innovations Association FRUCT

ISSN: 2305-7254 (Print); 2343-0737 (Online)
Publisher: FRUCT
Country of publisher: Finland
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering: Telecommunication
Website: http://fruct.org/publication

About the journal

Abstract

Keywords