KFSENet: A Key Frame-Based Skeleton Feature Estimation and Action Recognition Network for Improved Robot Vision with Face and Emotion Recognition

Dinh-Son Le; Hai-Hong Phan; Ha Huy Hung; Van-An Tran; The-Hung Nguyen; Dinh-Quan Nguyen

doi:10.3390/app12115455

Applied Sciences (May 2022)

KFSENet: A Key Frame-Based Skeleton Feature Estimation and Action Recognition Network for Improved Robot Vision with Face and Emotion Recognition

Dinh-Son Le,
Hai-Hong Phan,
Ha Huy Hung,
Van-An Tran,
The-Hung Nguyen,
Dinh-Quan Nguyen

Affiliations

Dinh-Son Le: Faculty of Information Technology, Le Quy Don Technical University, 236 Hoang Quoc Viet, Bac Tu Liem, Ha Noi 11900, Vietnam
Hai-Hong Phan: Faculty of Information Technology, Le Quy Don Technical University, 236 Hoang Quoc Viet, Bac Tu Liem, Ha Noi 11900, Vietnam
Ha Huy Hung: Faculty of Aerospace Engineering, Le Quy Don Technical University, 236 Hoang Quoc Viet, Bac Tu Liem, Ha Noi 11900, Vietnam
Van-An Tran: Faculty of Information Technology, Le Quy Don Technical University, 236 Hoang Quoc Viet, Bac Tu Liem, Ha Noi 11900, Vietnam
The-Hung Nguyen: Faculty of Technical Management, Le Quy Don Technical University, 236 Hoang Quoc Viet, Bac Tu Liem, Ha Noi 11900, Vietnam
Dinh-Quan Nguyen: Faculty of Aerospace Engineering, Le Quy Don Technical University, 236 Hoang Quoc Viet, Bac Tu Liem, Ha Noi 11900, Vietnam

DOI: https://doi.org/10.3390/app12115455
Journal volume & issue: Vol. 12, no. 11
p. 5455

Abstract

Read online

In this paper, we propose an integrated approach to robot vision: a key frame-based skeleton feature estimation and action recognition network (KFSENet) that incorporates action recognition with face and emotion recognition to enable social robots to engage in more personal interactions. Instead of extracting the human skeleton features from the entire video, we propose a key frame-based approach for their extraction using pose estimation models. We select the key frames using the gradient of a proposed total motion metric that is computed using dense optical flow. We use the extracted human skeleton features from the selected key frames to train a deep neural network (i.e., the double-feature double-motion network (DDNet)) for action recognition. The proposed KFSENet utilizes a simpler model to learn and differentiate between the different action classes, is computationally simpler and yields better action recognition performance when compared with existing methods. The use of key frames allows the proposed method to eliminate unnecessary and redundant information, which improves its classification accuracy and decreases its computational cost. The proposed method is tested on both publicly available standard benchmark datasets and self-collected datasets. The performance of the proposed method is compared to existing state-of-the-art methods. Our results indicate that the proposed method yields better performance compared with existing methods. Moreover, our proposed framework integrates face and emotion recognition to enable social robots to engage in more personal interaction with humans.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords