Long-Term Visitation Value for Deep Exploration in Sparse-Reward Reinforcement Learning

Simone Parisi; Davide Tateo; Maximilian Hensel; Carlo D’Eramo; Jan Peters; Joni Pajarinen

doi:10.3390/a15030081

Algorithms (Feb 2022)

Long-Term Visitation Value for Deep Exploration in Sparse-Reward Reinforcement Learning

Simone Parisi,
Davide Tateo,
Maximilian Hensel,
Carlo D’Eramo,
Jan Peters,
Joni Pajarinen

Affiliations

Simone Parisi: Meta AI Research, 4720 Forbes Avenue, Pittsburgh, PA 15213, USA
Davide Tateo: Department of Informatics, Technische Universität Darmstadt, Hochschulstr. 10, 64289 Darmstadt, Germany
Maximilian Hensel: Department of Informatics, Technische Universität Darmstadt, Hochschulstr. 10, 64289 Darmstadt, Germany
Carlo D’Eramo: Department of Informatics, Technische Universität Darmstadt, Hochschulstr. 10, 64289 Darmstadt, Germany
Jan Peters: Department of Informatics, Technische Universität Darmstadt, Hochschulstr. 10, 64289 Darmstadt, Germany
Joni Pajarinen: Department of Electrical Engineering and Automation, Aalto University, Tietotekniikantalo, Konemiehentie 2, 02150 Espoo, Finland

DOI: https://doi.org/10.3390/a15030081
Journal volume & issue: Vol. 15, no. 3
p. 81

Abstract

Read online

Reinforcement learning with sparse rewards is still an open challenge. Classic methods rely on getting feedback via extrinsic rewards to train the agent, and in situations where this occurs very rarely the agent learns slowly or cannot learn at all. Similarly, if the agent receives also rewards that create suboptimal modes of the objective function, it will likely prematurely stop exploring. More recent methods add auxiliary intrinsic rewards to encourage exploration. However, auxiliary rewards lead to a non-stationary target for the Q-function. In this paper, we present a novel approach that (1) plans exploration actions far into the future by using a long-term visitation count, and (2) decouples exploration and exploitation by learning a separate function assessing the exploration value of the actions. Contrary to existing methods that use models of reward and dynamics, our approach is off-policy and model-free. We further propose new tabular environments for benchmarking exploration in reinforcement learning. Empirical results on classic and novel benchmarks show that the proposed approach outperforms existing methods in environments with sparse rewards, especially in the presence of rewards that create suboptimal modes of the objective function. Results also suggest that our approach scales gracefully with the size of the environment.

Published in Algorithms

ISSN: 1999-4893 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Technology (General): Industrial engineering. Management engineering; Science: Mathematics: Instruments and machines: Electronic computers. Computer science
Website: https://www.mdpi.com/journal/algorithms

About the journal

Abstract

Keywords