Chinese image captioning with fusion encoder and visual keyword search

Yang Zou; Shiyu Liao; Qifei Wang

doi:10.1049/ipr2.13155

IET Image Processing (Sep 2024)

Chinese image captioning with fusion encoder and visual keyword search

Yang Zou,
Shiyu Liao,
Qifei Wang

Affiliations

Yang Zou: Institute of Intelligence Science and Technology College of Computer and Information Hohai University Nanjing China
Shiyu Liao: Institute of Intelligence Science and Technology College of Computer and Information Hohai University Nanjing China
Qifei Wang: Institute of Intelligence Science and Technology College of Computer and Information Hohai University Nanjing China

DOI: https://doi.org/10.1049/ipr2.13155
Journal volume & issue: Vol. 18, no. 11
pp. 3055 – 3069

Abstract

Read online

Abstract Automatic generation of image captions is essentially a cross‐modal conversion from image to text. Owing to the differences in linguistic characteristics between Chinese and English, quite a few Chinese image captioning methods have recently been proposed. Nevertheless, the existing Chinese image captioning models usually lack attention to local details of images or tend to produce general descriptions. To address these challenges, a Chinese image captioning method is proposed that incorporates fusion encoder, visual keyword search, and reinforcement learning. The fusion encoder can simultaneously extract local and global features of the input image to enrich the semantic information in the decoding stage, visual keyword search can pursue potential visual words associated with the image content, and the reinforcement learning mechanism can optimize the evaluation metric CIDEr at sentence level to promote the lexical diversity of image description. The results of extensive experiments demonstrate that the proposed model outperforms the state‐of‐the‐art models and delivers expressive and informative Chinese image captions.

Published in IET Image Processing

ISSN: 1751-9659 (Print); 1751-9667 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Technology: Photography; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/17519667

About the journal

Abstract

Keywords