MILG: Realistic lip-sync video generation with audio-modulated image inpainting

Han Bao; Xuhong Zhang; Qinying Wang; Kangming Liang; Zonghui Wang; Shouling Ji; Wenzhi Chen

Visual Informatics (Sep 2024)

MILG: Realistic lip-sync video generation with audio-modulated image inpainting

Han Bao,
Xuhong Zhang,
Qinying Wang,
Kangming Liang,
Zonghui Wang,
Shouling Ji,
Wenzhi Chen

Affiliations

Han Bao: School of Software Technology, Zhejiang University, Hangzhou, China
Xuhong Zhang: School of Software Technology, Zhejiang University, Hangzhou, China
Qinying Wang: College of Computer Science, Zhejiang University, Hangzhou, China
Kangming Liang: College of Engineering, Zhejiang University, Hangzhou, China
Zonghui Wang: College of Computer Science, Zhejiang University, Hangzhou, China; Corresponding author.
Shouling Ji: College of Computer Science, Zhejiang University, Hangzhou, China
Wenzhi Chen: College of Computer Science, Zhejiang University, Hangzhou, China

Journal volume & issue: Vol. 8, no. 3
pp. 71 – 81

Abstract

Read online

Existing lip synchronization (lip-sync) methods generate accurately synchronized mouths and faces in a generated video. However, they still confront the problem of artifacts in regions of non-interest (RONI), e.g., background and other parts of a face, which decreases the overall visual quality. To solve these problems, we innovatively introduce diverse image inpainting to lip-sync generation. We propose Modulated Inpainting Lip-sync GAN (MILG), an audio-constraint inpainting network to predict synchronous mouths. MILG utilizes prior knowledge of RONI and audio sequences to predict lip shape instead of image generation, which can keep the RONI consistent. Specifically, we integrate modulated spatially probabilistic diversity normalization (MSPD Norm) in our inpainting network, which helps the network generate fine-grained diverse mouth movements guided by the continuous audio features. Furthermore, to lower the training overhead, we modify the contrastive loss in lip-sync to support small-batch-size and few-sample training. Extensive experiments demonstrate that our approach outperforms the existing state-of-the-art of image quality and authenticity while keeping lip-sync.

Published in Visual Informatics

ISSN: 2468-502X (Online)
Publisher: Elsevier
Country of publisher: Netherlands
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: https://www.journals.elsevier.com/visual-informatics/

About the journal

Abstract

Keywords