NATCA YOLO-Based Small Object Detection for Aerial Images

Yicheng Zhu; Zhenhua Ai; Jinqiang Yan; Silong Li; Guowei Yang; Teng Yu

doi:10.3390/info15070414

Information (Jul 2024)

NATCA YOLO-Based Small Object Detection for Aerial Images

Yicheng Zhu,
Zhenhua Ai,
Jinqiang Yan,
Silong Li,
Guowei Yang,
Teng Yu

Affiliations

Yicheng Zhu: College of Electronic Information, Qingdao University, Qingdao 266071, China
Zhenhua Ai: College of Electronic Information, Qingdao University, Qingdao 266071, China
Jinqiang Yan: College of Electronic Information, Qingdao University, Qingdao 266071, China
Silong Li: College of Electronic Information, Qingdao University, Qingdao 266071, China
Guowei Yang: College of Electronic Information, Qingdao University, Qingdao 266071, China
Teng Yu: College of Electronic Information, Qingdao University, Qingdao 266071, China

DOI: https://doi.org/10.3390/info15070414
Journal volume & issue: Vol. 15, no. 7
p. 414

Abstract

Read online

The object detection model in UAV aerial image scenes faces challenges such as significant scale changes of certain objects and the presence of complex backgrounds. This paper aims to address the detection of small objects in aerial images using NATCA (neighborhood attention Transformer coordinate attention) YOLO. Specifically, the feature extraction network incorporates a neighborhood attention transformer (NAT) into the last layer to capture global context information and extract diverse features. Additionally, the feature fusion network (Neck) incorporates a coordinate attention (CA) module to capture channel information and longer-range positional information. Furthermore, the activation function in the original convolutional block is replaced with Meta-ACON. The NAT serves as the prediction layer in the new network, which is evaluated using the VisDrone2019-DET object detection dataset as a benchmark, and tested on the VisDrone2019-DET-test-dev dataset. To assess the performance of the NATCA YOLO model in detecting small objects in aerial images, other detection networks, such as Faster R-CNN, RetinaNet, and SSD, are employed for comparison on the test set. The results demonstrate that the NATCA YOLO detection achieves an average accuracy of 42%, which is a 2.9% improvement compared to the state-of-the-art detection network TPH-YOLOv5.

Published in Information

ISSN: 2078-2489 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering: Information technology
Website: http://www.mdpi.com/journal/information/

About the journal

Abstract

Keywords