Edge‐guided representation learning for underwater object detection

Linhui Dai; Hong Liu; Pinhao Song; Hao Tang; Runwei Ding; Shengquan Li

doi:10.1049/cit2.12325

CAAI Transactions on Intelligence Technology (Oct 2024)

Edge‐guided representation learning for underwater object detection

Linhui Dai,
Hong Liu,
Pinhao Song,
Hao Tang,
Runwei Ding,
Shengquan Li

Affiliations

Linhui Dai: Key Laboratory of Machine Perception Shenzhen Graduate School Peking University Shenzhen China
Hong Liu: Key Laboratory of Machine Perception Shenzhen Graduate School Peking University Shenzhen China
Pinhao Song: Robotics Research Group KU Leuven Leuven Belgium
Hao Tang: Computer Vision Lab ETH Zurich Zurich Switzerland
Runwei Ding: Peng Cheng Laboratory Shenzhen China
Shengquan Li: Peng Cheng Laboratory Shenzhen China

DOI: https://doi.org/10.1049/cit2.12325
Journal volume & issue: Vol. 9, no. 5
pp. 1078 – 1091

Abstract

Read online

Abstract Underwater object detection (UOD) is crucial for marine economic development, environmental protection, and the planet's sustainable development. The main challenges of this task arise from low‐contrast, small objects, and mimicry of aquatic organisms. The key to addressing these challenges is to focus the model on obtaining more discriminative information. The authors observe that the edges of underwater objects are highly unique and can be distinguished from low‐contrast or mimicry environments based on their edges. Motivated by this observation, an Edge‐guided Representation Learning Network, termed ERL‐Net is proposed, that aims to achieve discriminative representation learning and aggregation under the guidance of edge cues. Firstly, an edge‐guided attention module is introduced to model the explicit boundary information, which generates more discriminative features. Secondly, a hierarchical feature aggregation module is proposed to aggregate the multi‐scale discriminative features by regrouping them into three levels, effectively aggregating global and local information for locating and recognising underwater objects. Finally, a wide and asymmetric receptive field block is proposed to enable features to have a wider receptive field, allowing the model to focus on smaller object information. Comprehensive experiments on three challenging underwater datasets show that our method achieves superior performance on the UOD task.

Published in CAAI Transactions on Intelligence Technology

ISSN: 2468-2322 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Language and Literature: Philology. Linguistics: Computational linguistics. Natural language processing; Science: Mathematics: Instruments and machines: Electronic computers. Computer science: Computer software
Website: https://ietresearch.onlinelibrary.wiley.com/journal/24682322

About the journal

Abstract

Keywords