Double Deep Q-Network with Dynamic Bootstrapping for Real-Time Isolated Signal Control: A Traffic Engineering Perspective

Qiming Zheng; Hongfeng Xu; Jingyun Chen; Dong Zhang; Kun Zhang; Guolei Tang

doi:10.3390/app12178641

Applied Sciences (Aug 2022)

Double Deep Q-Network with Dynamic Bootstrapping for Real-Time Isolated Signal Control: A Traffic Engineering Perspective

Qiming Zheng,
Hongfeng Xu,
Jingyun Chen,
Dong Zhang,
Kun Zhang,
Guolei Tang

Affiliations

Qiming Zheng: School of Transportation and Logistics, Dalian University of Technology, Dalian 116024, China
Hongfeng Xu: School of Transportation and Logistics, Dalian University of Technology, Dalian 116024, China
Jingyun Chen: School of Transportation and Logistics, Dalian University of Technology, Dalian 116024, China
Dong Zhang: School of Transportation and Logistics, Dalian University of Technology, Dalian 116024, China
Kun Zhang: School of Transportation and Logistics, Dalian University of Technology, Dalian 116024, China
Guolei Tang: School of Port, Waterway and Ocean Engineering, Dalian University of Technology, Dalian 116024, China

DOI: https://doi.org/10.3390/app12178641
Journal volume & issue: Vol. 12, no. 17
p. 8641

Abstract

Read online

Real-time isolated signal control (RISC) at an intersection is of interest in the field of traffic engineering. Energizing RISC with reinforcement learning (RL) is feasible and necessary. Previous studies paid less attention to traffic engineering considerations and under-utilized traffic expertise to construct RL tasks. This study profiles the single-ring RISC problem from the perspective of traffic engineers, and improves a prevailing RL method for solving it. By qualitative applicability analysis, we choose double deep Q-network (DDQN) as the basic method. A single agent is deployed for an intersection. Reward is defined with vehicle departures to properly encourage and punish the agent’s behavior. The action is to determine the remaining green time for the current vehicle phase. State is represented in a grid-based mode. To update action values in time-varying environments, we present a temporal-difference algorithm TD(Dyn) to perform dynamic bootstrapping with the variable interval between actions selected. To accelerate training, we propose a data augmentation based on intersection symmetry. Our improved DDQN, termed D3ynQN, is subject to the signal timing constraints in engineering. The experiments at a close-to-reality intersection indicate that, by means of D3ynQN and non-delay-based reward, the agent acquires useful knowledge to significantly outperform a fully-actuated control technique in reducing average vehicle delay.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords