Multi-Subject Image Retrieval by Fusing Object and Scene-Level Feature Embeddings

Chung-Gi Ban; Youngbae Hwang; Dayoung Park; Ryong Lee; Rae-Young Jang; Myung-Seok Choi

doi:10.3390/app122412705

Applied Sciences (Dec 2022)

Multi-Subject Image Retrieval by Fusing Object and Scene-Level Feature Embeddings

Chung-Gi Ban,
Youngbae Hwang,
Dayoung Park,
Ryong Lee,
Rae-Young Jang,
Myung-Seok Choi

Affiliations

Chung-Gi Ban: Department of Control and Robot Engineering, Chungbuk National University, Cheongju 28644, Republic of Korea
Youngbae Hwang: Department of Control and Robot Engineering, Chungbuk National University, Cheongju 28644, Republic of Korea
Dayoung Park: Department of Control and Robot Engineering, Chungbuk National University, Cheongju 28644, Republic of Korea
Ryong Lee: Department of Machine Learning Data Research, Korea Institute of Science and Technology Information (KISTI), Daejeon 34141, Republic of Korea
Rae-Young Jang: Department of Machine Learning Data Research, Korea Institute of Science and Technology Information (KISTI), Daejeon 34141, Republic of Korea
Myung-Seok Choi: Department of Machine Learning Data Research, Korea Institute of Science and Technology Information (KISTI), Daejeon 34141, Republic of Korea

DOI: https://doi.org/10.3390/app122412705
Journal volume & issue: Vol. 12, no. 24
p. 12705

Abstract

Read online

Most existing image retrieval methods separately retrieve single images, such as a scene, content, or object, from a single database. However, for general purposes, target databases for image retrieval can include multiple subjects because it is not easy to predict which subject is entered. In this paper, we propose that image retrieval can be performed in practical applications by combining multiple databases. To deal with multi-subject image retrieval (MSIR), image embedding is generated through the fusion of scene- and object-level features, which are based on Detection Transformer (DETR) and a random patch generator with a deep-learning network, respectively. To utilize these feature vectors for image retrieval, two bags-of-visual-words (BoVWs) were used as feature embeddings because they are simply integrated with preservation of the characteristics of both features. A fusion strategy between the two BoVWs was proposed in three stages. Experiments were conducted to compare the proposed method with previous methods on conventional single-subject datasets and multi-subject datasets. The results validated that the proposed fused feature embeddings are effective for MSIR.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords