Dynamic deformable transformer for end‐to‐end face alignment

Liming Han; Chi Yang; Qing Li; Bin Yao; Zixian Jiao; Qianyang Xie

doi:10.1049/cvi2.12208

IET Computer Vision (Dec 2023)

Dynamic deformable transformer for end‐to‐end face alignment

Liming Han,
Chi Yang,
Qing Li,
Bin Yao,
Zixian Jiao,
Qianyang Xie

Affiliations

Liming Han: Institute of Microelectronics of the Chinese Academy of Sciences University of Chinese Academy of Sciences Beijing China
Chi Yang: Department of Oral Surgery, Shanghai Ninth People’s Hospital Shanghai Jiao Tong University School of Medicine Shanghai China
Qing Li: Institute of Microelectronics of the Chinese Academy of Sciences University of Chinese Academy of Sciences Beijing China
Bin Yao: Institute of Microelectronics of the Chinese Academy of Sciences University of Chinese Academy of Sciences Beijing China
Zixian Jiao: Department of Oral Surgery, Shanghai Ninth People’s Hospital Shanghai Jiao Tong University School of Medicine Shanghai China
Qianyang Xie: Department of Oral Surgery, Shanghai Ninth People’s Hospital Shanghai Jiao Tong University School of Medicine Shanghai China

DOI: https://doi.org/10.1049/cvi2.12208
Journal volume & issue: Vol. 17, no. 8
pp. 948 – 961

Abstract

Read online

Abstract Heatmap‐based regression (HBR) methods have dominated for a long time in the face alignment field while these methods need complex design and post‐processing. In this study, the authors propose an end‐to‐end and simple enough coordinate‐based regression (CBR) method called Dynamic Deformable Transformer (DDT) for face alignment. Unlike general pre‐defined landmark queries, DDT uses Dynamic Landmark Queries (DLQs) to query landmarks' classes and coordinates together. Besides, DDT adopts a deformable attention mechanism rather than a regular attention mechanism which has a faster convergence speed and lower computational complexity. Experiment results on three mainstream datasets 300W, WFLW, and COFW demonstrate DDT exceeds the state‐of‐the‐art CBR methods by a large margin and is comparable to the current state‐of‐the‐art HBR methods with much less computational complexity.

Published in IET Computer Vision

ISSN: 1751-9632 (Print); 1751-9640 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Medicine: Medicine (General): Computer applications to medicine. Medical informatics; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/17519640

About the journal

Abstract

Keywords