Multi-View Attention Network for Visual Dialog

Sungjin Park; Taesun Whang; Yeochan Yoon; Heuiseok Lim

doi:10.3390/app11073009

Applied Sciences (Mar 2021)

Multi-View Attention Network for Visual Dialog

Sungjin Park,
Taesun Whang,
Yeochan Yoon,
Heuiseok Lim

Affiliations

Sungjin Park: Department of Computer Science and Engineering, Korea University, 145, Anam-ro, Seongbuk-gu, Seoul 02841, Korea
Taesun Whang: Wisenut Inc., 49, Daewangpangyo-ro 644beon-gil, Bundang-gu, Seongnam-si 13493, Gyeonggi-do, Korea
Yeochan Yoon: Electronics and Technology Research Institute (ETRI), 161, Gajeong-dong, Yuseong-gu, Daejeon 34129, Korea
Heuiseok Lim: Department of Computer Science and Engineering, Korea University, 145, Anam-ro, Seongbuk-gu, Seoul 02841, Korea

DOI: https://doi.org/10.3390/app11073009
Journal volume & issue: Vol. 11, no. 7
p. 3009

Abstract

Read online

Visual dialog is a challenging vision-language task in which a series of questions visually grounded by a given image are answered. To resolve the visual dialog task, a high-level understanding of various multimodal inputs (e.g., question, dialog history, and image) is required. Specifically, it is necessary for an agent to (1) determine the semantic intent of question and (2) align question-relevant textual and visual contents among heterogeneous modality inputs. In this paper, we propose Multi-View Attention Network (MVAN), which leverages multiple views about heterogeneous inputs based on attention mechanisms. MVAN effectively captures the question-relevant information from the dialog history with two complementary modules (i.e., Topic Aggregation and Context Matching), and builds multimodal representations through sequential alignment processes (i.e., Modality Alignment). Experimental results on VisDial v1.0 dataset show the effectiveness of our proposed model, which outperforms previous state-of-the-art methods under both single model and ensemble settings.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords