GITPose: going shallow and deeper using vision transformers for human pose estimation

Evans Aidoo; Xun Wang; Zhenguang Liu; Abraham Opanfo Abbam; Edwin Kwadwo Tenagyei; Victor Nonso Ejianya; Seth Larweh Kodjiku; Esther Stacy E. B. Aggrey

doi:10.1007/s40747-024-01361-y

Complex & Intelligent Systems (Mar 2024)

GITPose: going shallow and deeper using vision transformers for human pose estimation

Evans Aidoo,
Xun Wang,
Zhenguang Liu,
Abraham Opanfo Abbam,
Edwin Kwadwo Tenagyei,
Victor Nonso Ejianya,
Seth Larweh Kodjiku,
Esther Stacy E. B. Aggrey

Affiliations

Evans Aidoo: School of statistics and mathematics, Zhejiang Gongsheng University
Xun Wang: School of Computer Science and Technology, Zhejiang Gongshang University
Zhenguang Liu: School of Cyber Science and Technology, Zhejiang University
Abraham Opanfo Abbam: School of Computer Science and Technology, Zhejiang Gongshang University
Edwin Kwadwo Tenagyei: School of Engineering and Built Environment, Griffith University
Victor Nonso Ejianya: School of statistics and mathematics, Zhejiang Gongsheng University
Seth Larweh Kodjiku: School of statistics and mathematics, Zhejiang Gongsheng University
Esther Stacy E. B. Aggrey: School of Information and Software Engineering, University of Electronic Science and Technology of China

DOI: https://doi.org/10.1007/s40747-024-01361-y
Journal volume & issue: Vol. 10, no. 3
pp. 4507 – 4520

Abstract

Read online

Abstract In comparison to convolutional neural networks (CNN), the newly created vision transformer (ViT) has demonstrated impressive outcomes in human pose estimation (HPE). However, (1) there is a quadratic rise in complexity with respect to image size, which causes the traditional ViT to be unsuitable for scaling, and (2) the attention process at the transformer encoder as well as decoder also adds substantial computational costs to the detector’s overall processing time. Motivated by this, we propose a novel Going shallow and deeper with vIsion Transformers for human Pose estimation (GITPose) without CNN backbones for feature extraction. In particular, we introduce a hierarchical transformer in which we utilize multilayer perceptrons to encode the richest local feature tokens in the initial phases (i.e., shallow), whereas self-attention modules are employed to encode long-term relationships in the deeper layers (i.e., deeper), and a decoder for keypoint detection. In addition, we offer a learnable deformable token association module (DTA) to non-uniformly and dynamically combine informative keypoint tokens. Comprehensive evaluation and testing on the COCO and MPII benchmark datasets reveal that GITPose achieves a competitive average precision (AP) on pose estimation compared to its state-of-the-art approaches.

Published in Complex & Intelligent Systems

ISSN: 2199-4536 (Print); 2198-6053 (Online)
Publisher: Springer
Country of publisher: Switzerland
LCC subjects: Science: Mathematics: Instruments and machines: Electronic computers. Computer science; Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: https://www.springer.com/journal/40747

About the journal

Abstract

Keywords