Deep Reinforcement-Learning-Based Air-Combat-Maneuver Generation Framework

Junru Mei; Ge Li; Hesong Huang

doi:10.3390/math12193020

Mathematics (Sep 2024)

Deep Reinforcement-Learning-Based Air-Combat-Maneuver Generation Framework

Junru Mei,
Ge Li,
Hesong Huang

Affiliations

Junru Mei: College of Systems Engineering, National University of Defense Technology, Changsha 410073, China
Ge Li: College of Systems Engineering, National University of Defense Technology, Changsha 410073, China
Hesong Huang: College of Systems Engineering, National University of Defense Technology, Changsha 410073, China

DOI: https://doi.org/10.3390/math12193020
Journal volume & issue: Vol. 12, no. 19
p. 3020

Abstract

Read online

With the development of unmanned aircraft and artificial intelligence technology, the future of air combat is moving towards unmanned and autonomous direction. In this paper, we introduce a new layered decision framework designed to address the six-degrees-of-freedom (6-DOF) aircraft within-visual-range (WVR) air-combat challenge. The decision-making process is divided into two layers, each of which is addressed separately using reinforcement learning (RL). The upper layer is the combat policy, which determines maneuvering instructions based on the current combat situation (such as altitude, speed, and attitude). The lower layer control policy then uses these commands to calculate the input signals from various parts of the aircraft (aileron, elevator, rudder, and throttle). Among them, the control policy is modeled as a Markov decision framework, and the combat policy is modeled as a partially observable Markov decision framework. We describe the two-layer training method in detail. For the control policy, we designed rewards based on expert knowledge to accurately and stably complete autonomous driving tasks. At the same time, for combat policy, we introduce a self-game-based course learning, allowing the agent to play against historical policies during training to improve performance. The experimental results show that the operational success rate of the proposed method against the game theory baseline reaches 85.7%. Efficiency was also outstanding, with an average 13.6% reduction in training time compared to the RL baseline.

Published in Mathematics

ISSN: 2227-7390 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Science: Mathematics
Website: http://www.mdpi.com/journal/mathematics

About the journal

Abstract

Keywords