Query by Example: Semantic Traffic Scene Retrieval Using LLM-Based Scene Graph Representation

Yafu Tian; Alexander Carballo; Ruifeng Li; Simon Thompson; Kazuya Takeda

doi:10.3390/s25082546

Sensors (Apr 2025)

Query by Example: Semantic Traffic Scene Retrieval Using LLM-Based Scene Graph Representation

Yafu Tian,
Alexander Carballo,
Ruifeng Li,
Simon Thompson,
Kazuya Takeda

Affiliations

Yafu Tian: Graduate School of Informatics, Nagoya University, Furo-cho, Chikusa-ku, Nagoya 464-8603, Japan
Alexander Carballo: Tier IV Inc., Nagoya University Open Innovation Center, 1-3, Mei-eki 1-chome, Nakamura-Ward, Nagoya 450-6610, Japan
Ruifeng Li: State Key Laboratory of Robotic and Intelligent System, Harbin Institute of Technology, Harbin 150000, China
Simon Thompson: Tier IV Inc., Nagoya University Open Innovation Center, 1-3, Mei-eki 1-chome, Nakamura-Ward, Nagoya 450-6610, Japan
Kazuya Takeda: Graduate School of Informatics, Nagoya University, Furo-cho, Chikusa-ku, Nagoya 464-8603, Japan

DOI: https://doi.org/10.3390/s25082546
Journal volume & issue: Vol. 25, no. 8
p. 2546

Abstract

Read online

In autonomous driving, retrieving a specific traffic scene in huge datasets is a significant challenge. Traditional scene retrieval methods struggle to cope with the semantic complexity and heterogeneity of traffic scenes and are unable to meet the variable needs of different users. This paper proposes “Query-by-Example”, a traffic scene retrieval approach based on Visual-Large Language Model (VLM)-generated Road Scene Graph (RSG) representation. Our method uses VLMs to generate structured scene graphs from video data, capturing high-level semantic attributes and detailed object relationships in traffic scenes. We introduce an extensible set of scene attributes and a graph-based scene description to quantify scene similarity. We also propose a RSG-LLM benchmark dataset containing 1000 traffic scenes, their corresponding natural language descriptions, and RSGs to evaluate the performance of LLMs in generating RSGs. Experiments show that our method can effectively retrieve semantically similar traffic scenes from large databases, supporting various query formats, including natural language, images, video clips, rosbag, etc. Our method provides a comprehensive and flexible framework for traffic scene retrieval, promoting its application in autonomous driving systems.

Published in Sensors

ISSN: 1424-8220 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Chemical technology
Website: http://www.mdpi.com/journal/sensors

About the journal

Abstract

Keywords