A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations

Kexv Li; Yue Wang; Xing Zhuang; Hao Yin; Xinyu Liu; Hanyu Li

doi:10.3390/drones7040232

Drones (Mar 2023)

A Penetration Method for UAV Based on Distributed Reinforcement Learning and Demonstrations

Kexv Li,
Yue Wang,
Xing Zhuang,
Hao Yin,
Xinyu Liu,
Hanyu Li

Affiliations

Kexv Li: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing 100081, China
Yue Wang: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing 100081, China
Xing Zhuang: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing 100081, China
Hao Yin: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing 100081, China
Xinyu Liu: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing 100081, China
Hanyu Li: School of Mechatronical Engineering, Beijing Institute of Technology, Beijing 100081, China

DOI: https://doi.org/10.3390/drones7040232
Journal volume & issue: Vol. 7, no. 4
p. 232

Abstract

Read online

The penetration of unmanned aerial vehicles (UAVs) is an essential and important link in modern warfare. Enhancing UAV’s ability of autonomous penetration through machine learning has become a research hotspot. However, the current generation of autonomous penetration strategies for UAVs faces the problem of excessive sample demand. To reduce the sample demand, this paper proposes a combination policy learning (CPL) algorithm that combines distributed reinforcement learning and demonstrations. Innovatively, the action of the CPL algorithm is jointly determined by the initial policy obtained from demonstrations and the target policy in the asynchronous advantage actor-critic network, thus retaining the guiding role of demonstrations in the initial training. In a complex and unknown dynamic environment, 1000 training experiments and 500 test experiments were conducted for the CPL algorithm and related baseline algorithms. The results show that the CPL algorithm has the smallest sample demand, the highest convergence efficiency, and the highest success rate of penetration among all the algorithms, and has strong robustness in dynamic environments.

Published in Drones

ISSN: 2504-446X (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Motor vehicles. Aeronautics. Astronautics
Website: http://www.mdpi.com/journal/drones

About the journal

Abstract

Keywords