The magmaOffenburg 2020 RoboCup 3D Simulation Team
Nico Bohlinger, Hannes Braun, Klaus Dorer, Jens Fischer, Carmen Schmider, Jannik Seiler, David Weiler
Hochschule Offenburg, Elektrotechnik, Medizintechnik und Informatik, Germany
Abstract Team description papers of magmaOffenburg are incremental in the sense that each year we address a different topic of our team and the tools around our team. In this year's team description paper we address our approach to learn a model free walk with Nao toe using genetic algorithms.
1 Introduction
Utilizing toes of a humanoid robot is difficult for various reasons, one of which is that inverse kinematics is overdetermined with the introduction of toe joints. Nevertheless, a number of robots with either passive [1, 2] or active toe joints [3–5] have been developed. Recent work shows considerable progress on learning model-free behaviors using genetic learning [6] for kicking with toes and deep reinforcement learning [7, 8] for walking without toe joints. In this work we show that toe joints can significantly improve the walking behavior of a simulated Nao robot and can be learned model-free.
2 Approach
Learning is performed in the SimSpark simulator using a 24 DoF simulated Nao robot shown in Figure 1. Head joints and, for most experiments, arm joints have been excluded from learning. For an initial proof of concept, we used only three of the seven leg joints per leg to keep the dimensionality low, but as one might expect, the results of learning runs with the full seven produced much faster and more dynamic walks (though both resulted in working walk behaviors).
The fitness function subtracts a penalty for falling from the walked distance in x-direction in meters. There is also a penalty for the maximum deviation in y-direction reached during an episode, weighted by a factor. In practice, the values chosen for fallenPenalty and factor were usually 3 and 2 respectively.
$$fitness = distanceX - fallenPenalty - (maxY * factor)$$
Genes encode angles and angular speeds for each joint over four to eight keyframes, as well as the duration between keyframes. In case of four keyframes, this results in 180 parameters to learn if arms are included. With these parameters, the robot learns a single step and mirrors the movement to get a double step.
3 Results
Experiments were run with 200 individuals over 200 generations with 10 oversampling runs per robot to average out non-determinism. The overall runtime of each such learning run is 2.5 days on our hardware.
Figure 2 shows a sequence of images for a learned step. The behavior reaches a speed of 1.3 m/s compared to the 1.0 m/s of our model-based walk and 0.96 m/s for a walk behavior learned on the Nao robot without toes. The learned walk with toes is less stable, however, and shows a fall rate of 30% compared to 2% of the model-based walk.
References
- Ogura, Y., Shimomura, K., Kondo, H., Morishima, A., Okubo, T., Momoki, S., Lim, H., Takanishi, A.: Human-like Walking with Knee Stretched, Heel-contact and Toe- off Motion by a Humanoid Robot. Proceedings of the IEEE/RSJ International Conference on Intelligent Robots Systems (IROS) (2006) 3976–81
- Sellaouti, R., Stasse, O., Kajita, S., Yokoi, K., Kheddar, A.: Faster and Smoother Walking of Humanoid HRP-2 with Passive Toe Joints. Proceedings of the IEEE/RSJ International Conference on Intelligent Robots Systems (IROS) (2006)
- Buschmann, T., Lohmeier, S., Ulbrich, H.: Humanoid robot lola: Design and walking control. Journal of physiology Vol 103, Paris (2009) 141–148
- Tajima, R., Honda, D., Suga, K.: Fast running experiments involving a humanoid robot. (05 2009) 1571–1576
- Behnke, S.: Human-like walking using toes joint and straight stance leg. Proceedings of 3rd International Symposium on Adaptive Motion in Animals and Machines (AMAM) (11 2005)
- Dorer, K.: Learning to use toes in a humanoid robot. In Akiyama, H., Obst, O., Sammut, C., Tonidandel, F., eds.: RoboCup 2017: Robot World Cup XXI, Springer (1 2018) 168–179
- Heess, N., TB, D., Sriram, S., Lemmon, J., Merel, J., Wayne, G., Tassa, Y., Erez, T., Wang, Z., Eslami, S.M.A., Riedmiller, M.A., Silver, D.: Emergence of locomotion behaviours in rich environments. arXiv:1707.02286 (2017)
- Abreu, M., Lau, N., Sousa, A., Reis, L.P.: Learning to run faster in a humanoid robot soccer environment through reinforcement learning. RoboCup 2019: Robot World Cup XXIII (2019)