Off‐policy correction algorithm for double Q network based on deep reinforcement learning

Qingbo Zhang; Manlu Liu; Heng Wang; Weimin Qian; Xinglang Zhang

doi:10.1049/csy2.12102

IET Cyber-systems and Robotics (Dec 2023)

Off‐policy correction algorithm for double Q network based on deep reinforcement learning

Qingbo Zhang,
Manlu Liu,
Heng Wang,
Weimin Qian,
Xinglang Zhang

Affiliations

Qingbo Zhang: School of Information Engineering Southwest University of Science and Technology Mianyang China
Manlu Liu: School of Information Engineering Southwest University of Science and Technology Mianyang China
Heng Wang: School of Information Engineering Southwest University of Science and Technology Mianyang China
Weimin Qian: School of Information Engineering Southwest University of Science and Technology Mianyang China
Xinglang Zhang: School of Information Engineering Southwest University of Science and Technology Mianyang China

DOI: https://doi.org/10.1049/csy2.12102
Journal volume & issue: Vol. 5, no. 4
pp. n/a – n/a

Abstract

Read online

Abstract A deep reinforcement learning (DRL) method based on the deep deterministic policy gradient (DDPG) algorithm is proposed to address the problems of a mismatch between the needed training samples and the actual training samples during the training of intelligence, the overestimation and underestimation of the existence of Q‐values, and the insufficient dynamism of the intelligence policy exploration. This method introduces the Actor‐Critic Off‐Policy Correction (AC‐Off‐POC) reinforcement learning framework and an improved double Q‐value learning method, which enables the value function network in the target task to provide a more accurate evaluation of the policy network and converge to the optimal policy more quickly and stably to obtain higher value returns. The method is applied to multiple MuJoCo tasks on the Open AI Gym simulation platform. The experimental results show that it is better than the DDPG algorithm based solely on the different policy correction framework (AC‐Off‐POC) and the conventional DRL algorithm. The value of returns and stability of the double‐Q‐network off‐policy correction algorithm for the deep deterministic policy gradient (DCAOP‐DDPG) proposed by the authors are significantly higher than those of other DRL algorithms.

Published in IET Cyber-systems and Robotics

ISSN: 2631-6315 (Online)
Publisher: Wiley
Country of publisher: United Kingdom
LCC subjects: Science: Science (General): Cybernetics; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://ietresearch.onlinelibrary.wiley.com/journal/26316315

About the journal

Abstract

Keywords