Towards a Broad-Persistent Advising Approach for Deep Interactive Reinforcement Learning in Robotic Environments

Hung Son Nguyen; Francisco Cruz; Richard Dazeley

doi:10.3390/s23052681

Sensors (Mar 2023)

Towards a Broad-Persistent Advising Approach for Deep Interactive Reinforcement Learning in Robotic Environments

Hung Son Nguyen,
Francisco Cruz,
Richard Dazeley

Affiliations

Hung Son Nguyen: School of Information Technology, Deakin University, Geelong 3220, Australia
Francisco Cruz: School of Computer Science and Engineering, University of New South Wales, Sydney 2052, Australia
Richard Dazeley: School of Information Technology, Deakin University, Geelong 3220, Australia

DOI: https://doi.org/10.3390/s23052681
Journal volume & issue: Vol. 23, no. 5
p. 2681

Abstract

Read online

Deep Reinforcement Learning (DeepRL) methods have been widely used in robotics to learn about the environment and acquire behaviours autonomously. Deep Interactive Reinforcement 2 Learning (DeepIRL) includes interactive feedback from an external trainer or expert giving advice to help learners choose actions to speed up the learning process. However, current research has been limited to interactions that offer actionable advice to only the current state of the agent. Additionally, the information is discarded by the agent after a single use, which causes a duplicate process at the same state for a revisit. In this paper, we present Broad-Persistent Advising (BPA), an approach that retains and reuses the processed information. It not only helps trainers give more general advice relevant to similar states instead of only the current state, but also allows the agent to speed up the learning process. We tested the proposed approach in two continuous robotic scenarios, namely a cart pole balancing task and a simulated robot navigation task. The results demonstrated that the agent’s learning speed increased, as evidenced by the rising reward points of up to 37%, while maintaining the number of interactions required for the trainer, in comparison to the DeepIRL approach.

Published in Sensors

ISSN: 1424-8220 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Chemical technology
Website: http://www.mdpi.com/journal/sensors

About the journal

Abstract

Keywords