Dynamic facial expression recognition with pseudo‐label guided multi‐modal pre‐training

Bing Yin; Shi Yin; Cong Liu; Yanyong Zhang; Changfeng Xi; Baocai Yin; Zhenhua Ling

doi:10.1049/cvi2.12217

IET Computer Vision (Feb 2024)

Dynamic facial expression recognition with pseudo‐label guided multi‐modal pre‐training

Bing Yin,
Shi Yin,
Cong Liu,
Yanyong Zhang,
Changfeng Xi,
Baocai Yin,
Zhenhua Ling

Affiliations

Bing Yin: School of Information Science and Technology University of Science and Technology of China Hefei City Anhui Province China
Shi Yin: iFLYTEK Research Hefei City Anhui Province China
Cong Liu: iFLYTEK Research Hefei City Anhui Province China
Yanyong Zhang: School of Computer Science and Technology University of Science and Technology of China Hefei City Anhui Province China
Changfeng Xi: iFLYTEK Research Hefei City Anhui Province China
Baocai Yin: iFLYTEK Research Hefei City Anhui Province China
Zhenhua Ling: School of Information Science and Technology University of Science and Technology of China Hefei City Anhui Province China

DOI: https://doi.org/10.1049/cvi2.12217
Journal volume & issue: Vol. 18, no. 1
pp. 33 – 45

Abstract

Read online

Abstract Due to the huge cost of manual annotations, the labelled data may not be sufficient to train a dynamic facial expression (DFR) recogniser with good performance. To address this, the authors propose a multi‐modal pre‐training method with a pseudo‐label guidance mechanism to make full use of unlabelled video data for learning informative representations of facial expressions. First, the authors build a pre‐training dataset of videos with aligned vision and audio modals. Second, the vision and audio feature encoders are trained through an instance discrimination strategy and a cross‐modal alignment strategy on the pre‐training data. Third, the vision feature encoder is extended as a dynamic expression recogniser and is fine‐tuned on the labelled training data. Fourth, the fine‐tuned expression recogniser is adopted to predict pseudo‐labels for the pre‐training data, and then start a new pre‐training phase with the guidance of pseudo‐labels to alleviate the long‐tail distribution problem and the instance‐class confliction. Fifth, since the representations learnt with the guidance of pseudo‐labels are more informative, a new fine‐tuning phase is added to further boost the generalisation performance on the DFR recognition task. Experimental results on the Dynamic Facial Expression in the Wild dataset demonstrate the superiority of the proposed method.

Published in IET Computer Vision

ISSN: 1751-9632 (Print); 1751-9640 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/17519640

About the journal

Abstract

Keywords