Spiking Neural Network Discovers Energy-Efficient Hexapod Motion in Deep Reinforcement Learning

Katsumi Naya; Kyo Kutsuzawa; Dai Owaki; Mitsuhiro Hayashibe

doi:10.1109/ACCESS.2021.3126311

IEEE Access (Jan 2021)

Spiking Neural Network Discovers Energy-Efficient Hexapod Motion in Deep Reinforcement Learning

Katsumi Naya,
Kyo Kutsuzawa,
Dai Owaki,
Mitsuhiro Hayashibe

Affiliations

Katsumi Naya: ORCiD; Department of Robotics, Graduate School of Engineering, Neuro-Robotics Laboratory, Tohoku University, Sendai, Japan
Kyo Kutsuzawa: ORCiD; Department of Robotics, Graduate School of Engineering, Neuro-Robotics Laboratory, Tohoku University, Sendai, Japan
Dai Owaki: ORCiD; Department of Robotics, Graduate School of Engineering, Neuro-Robotics Laboratory, Tohoku University, Sendai, Japan
Mitsuhiro Hayashibe: ORCiD; Department of Robotics, Graduate School of Engineering, Neuro-Robotics Laboratory, Tohoku University, Sendai, Japan

DOI: https://doi.org/10.1109/ACCESS.2021.3126311
Journal volume & issue: Vol. 9
pp. 150345 – 150354

Abstract

Read online

In Deep Reinforcement Learning (DRL) for robotics application, it is important to find energy-efficient motions. For this purpose, a standard method is to set an action penalty in the reward to find the optimal motion considering the energy expenditure. This method is widely used for the simplicity of implementation. However, since the reward is a linear sum, if the penalty is too large, the system will fall into local minima and no moving solution can be obtained. In contrast, if the penalty is too small, the effect may not be sufficient. Therefore, it is necessary to adjust the amount of the penalty so that the agent always moves dynamically, and the energy-saving effect is sufficient. Nevertheless, since adjusting the hyperparameters is computationally expensive, we need a learning method that is robust to the penalty setting problem. We investigated on the Spiking Neural Network (SNN), which has been attracting attention for its computational efficiency and neuromorphic architecture. We conducted gait experiments using a hexapod agent while varying the energy penalty settings in the simulation environment. By applying SNN to the conventional state-of-the-art DRL algorithms, we examined whether the agent could explore for an optimal gait with a larger penalty variation and obtain an energy-efficient gait verified with Cost of Transport (CoT), a metric of energy efficiency for gait. Soft Actor-Critic (SAC)+SNN resulted in a CoT of 1.64, Twin Delayed Deep Deterministic policy gradient (TD3)+SNN resulted in a CoT of 2.21, and Deep Deterministic policy gradient (DDPG)+SNN resulted in a CoT of 2.08 (1.91 for normal SAC, 2.38 for TD3, and 2.40 for DDPG). DRL combined with SNN succeeded in learning more energy efficient gait with lower CoT.

Published in IEEE Access

ISSN: 2169-3536 (Online)
Publisher: IEEE
Country of publisher: United States
LCC subjects: Technology: Electrical engineering. Electronics. Nuclear engineering
Website: https://ieeexplore.ieee.org/xpl/RecentIssue.jsp?punumber=6287639

About the journal

Abstract

Keywords