Reward Function and Configuration Parameters in Machine Learning of a Four-Legged Walking Robot

Arkadiusz Kubacki; Marcin Adamek; Piotr Baran

doi:10.3390/app131810298

Applied Sciences (Sep 2023)

Reward Function and Configuration Parameters in Machine Learning of a Four-Legged Walking Robot

Arkadiusz Kubacki,
Marcin Adamek,
Piotr Baran

Affiliations

Arkadiusz Kubacki: Institute of Mechanical Technology, Poznan University of Technology, ul. Piotrowo 3, 60-695 Poznan, Poland
Marcin Adamek: Institute of Mechanical Technology, Poznan University of Technology, ul. Piotrowo 3, 60-695 Poznan, Poland
Piotr Baran: Institute of Mechanical Technology, Poznan University of Technology, ul. Piotrowo 3, 60-695 Poznan, Poland

DOI: https://doi.org/10.3390/app131810298
Journal volume & issue: Vol. 13, no. 18
p. 10298

Abstract

Read online

In contemporary times, the use of walking robots is gaining increasing popularity and is prevalent in various industries. The ability to navigate challenging terrains is one of the advantages that they have over other types of robots, but they also require more intricate control mechanisms. One way to simplify this issue is to take advantage of artificial intelligence through reinforcement learning. The reward function is one of the conditions that governs how learning takes place, determining what actions the agent is willing to take based on the collected data. Another aspect to consider is the predetermined values contained in the configuration file, which describe the course of the training. The correct tuning of them is crucial for achieving satisfactory results in the teaching process. The initial phase of the investigation involved assessing the currently prevalent forms of kinematics for walking robots. Based on this evaluation, the most suitable design was selected. Subsequently, the Unity3D development environment was configured using an ML-Agents toolkit, which supports machine learning. During the experiment, the impacts of the values defined in the configuration file and the form of the reward function on the course of training were examined. Movement algorithms were developed for various modifications for learning to use artificial neural networks.

Published in Applied Sciences

ISSN: 2076-3417 (Online)
Publisher: MDPI AG
Country of publisher: Switzerland
LCC subjects: Technology: Engineering (General). Civil engineering (General); Science: Biology (General); Science: Physics; Science: Chemistry
Website: http://www.mdpi.com/journal/applsci

About the journal

Abstract

Keywords