FC Portugal 3D Simulation Team: Team Description Paper 2017
Luís Paulo Reis, Nuno Lau, Abbas Abdolmaleki, Nima Shafii, Rui Ferreira, Artur Pereira, David Simões
DETI/UA - Electronics, Telecommunications and Informatics Dep., University of Aveiro, Portugal; DSI/EEUM - School of Engineering, University of Minho, Portugal; IEETA - Institute of Electronics and Telematics Engineering of Aveiro, Portugal; LIACC - Artificial Intelligence and Computer Science Lab., University of Porto, Portugal
Abstract FC Portugal 3D team is developed upon the structure of our previous Simulation league 2D/3D teams and our standard platform league team. Our research concerning the robot low-level skills is focused on developing behaviors that may be applied on real robots with minimal adaptation using model-based approaches. Our research on high-level soccer coordination methodologies and team playing is mainly focused on the adaptation of previously developed methodologies from our 2D soccer teams to the 3D humanoid environment and on creating new coordination methodologies based on the previously developed ones. The research-oriented development of our team has been pushing it to be one of the most competitive over the years (World champion in 2000 and Coach Champion in 2002, European champion in 2000 and 2001, Coach 2nd place in 2003 and 2004, European champion in Rescue Simulation and Simulation 3D in 2006, World Champion in Simulation 3D in Bremen 2006 and European champion in 2007, 2012, 2013, 2014 and 2015). This paper describes some of the main innovations of our 3D simulation league team during the last years. New low-level behaviors have been developed for the simulated humanoid agent, which was based on the controlling of the dynamics of the behaviors, by using the principle of physical modeling. A new generic learning framework has also been developed, which is used on top of the low-level behaviors to tune the behavior parameters. This paper also includes general information related to the design of our agent architecture. The current research is focused on improving the current learning framework by developing new learning algorithms to optimize low-level skills performance, developing a new omnidirectional kick engine and integrating high-level coordination mechanisms. Very good results were already achieved, in previous years concerning the improvement of lowlevel skills, the use of high level coordination methodologies and the use of machine learning methodologies.
1. Introduction
FC Portugal was built upon the low-level skills research conducted during previous years. Although there is still space for improvement in FC Portugal low-level skills, we feel that we currently have a very performing set of these skills. Our research on developing low-level behaviors is mainly focused on approaches which can also be applied on the real robot with minimal adaptation. In this matter we developed low-level skills using model-based approaches, in which the stability of humanoid behavior is modeled using physical systems. Control the stability dynamics of a humanoid robot is still challenging and it is one of the main research directions of our team. In section 4, we will explain our approaches to develop robust and agile soccer low-level skills.
As another main research direction, we are also focused on the high-level decision and cooperation mechanisms of our agents. For RoboCup 3D soccer simulation competition that was based on spheres (from 2004 to 2006), the decisive factor (like in the 2D competition) was the high-level reasoning capacities of the players and not their low-level skills. Thus we worked mainly on high-level coordination methodologies for our previous teams. Since 2007 humanoid agents have been introduced in the 3D Simulation league, but the number of agents has been kept small until 2011. During this period research in coordination was not very important in the 3D league. Developing efficient low-level skills, contrarily to what should be the research focus of the simulation league, has been the main decisive factor in the 3D league, during this period. However, in 2011 the number of agents has increased to 9, and in 2012 teams were composed by 11 players making finally coordination, a very important issue for the efficiency of the team.
Our research on high-level soccer coordination methodologies and team playing is mainly focused on the adaptation of previously developed methodologies from our 2D soccer teams [1, 2, 3, 4, 5] to the 3D humanoid environment and on creating new coordination methodologies based on the previously developed ones. In our 2D teams, which participated in RoboCup since 2000 with very good results, we have introduced several concepts and algorithms covering a broad spectrum of the soccer simulation research challenges. From coordination techniques such as Tactics, Formations, Dynamic Positioning and Role Exchange, Situation Based Strategic Positioning and Intelligent Perception to Optimization based low-level skills, Visual Debugging and Coaching, the number of research aspects FC Portugal has been working on is quite extensive [1, 2, 3, 4, 5].
Several interesting topics were opened by the introduction of humanoid agents, including in the use of learning and optimization techniques for developing efficient both high-level and low-level skills. In previous work, we have introduced methods for developing very efficient low-level skills using optimization techniques [1, 6]. Recently, we have developed a new learning framework in which several optimization techniques have been included such as hill climbing (HC), tabu search (TS), genetic algorithms (GA), particle swarm optimization (PSO), Covariance Matrix Adaptation Evolution Strategy (CMA-ES) and other policy learning methods. This work has already conducted to the development of an efficient set of humanoid low-level skills. Section 6 presents briefly our new learning framework.
2. Research Directions
New research directions include research on developing our current layered architectures for agent controlling to be optimized and more efficient. Thus our research will be focused on improving both lower layers and higher layers. The lower layers will be responsible for the basic control of the humanoid such as stability, while the higher layers take decisions at a strategic level.
In our lower level control architecture, we will develop new kick skill based on our previous kick skill; also we will improve our developed running skill to be used as a primary locomotion for our robots. The robustness of the walking and running skill in face of the external forces will also be improved.
In the upper level control architecture, directions of research in FC Portugal include developing a model for a strategy for a humanoid game and the integration of humanoids coming from different teams in an inter-team framework to allow the formation of a team with different humanoids, and developing a new opponent modeling approach to model the opponent basic behaviors performance, its positioning, etc. These are factors that must be taken into account when selecting a given strategy for a game.
One of our improvements that have been achieved during the last year was implementing a new learning framework. This new learning framework guides us to optimize our robot behavior more efficiently. Several optimization and learning methods for generation of humanoid behaviors are being compared, including simulated annealing (SA), Hill Climbing (HC), GA, PSO, CMA-ES and PoWER. These techniques have been combined with physics based models and optimum control to derive very efficient skills.
Also heterogeneity will be important because in the future it is expected that not all humanoids will be identical, having humanoids with different capabilities introduces new problems of task assignment that will have to be dealt with in humanoid teams. We have already tested the use of heterogeneous humanoids in 2014 3D Simulation competition. We will improve our robots behavior by optimizing behavior specifications for each heterogeneous humanoid robot.
3. Agent Architecture
The FC Portugal Agent 3D [7] is divided in several packages: each one with a specific purpose. Figure 1 shows the general structure of the humanoid agent.
4. Low-Level Skills
In this section we briefly describe our approaches to develop soccer low-level skills, such as walking, running, and kicking. Nowadays, in order to compete well in RoboCup soccer humanoid leagues, the robots should be able to perform their low level skill fast, in an omnidirectional manner, and also robust against the external perturbation and noises. In order to improve our low-level skills, we model the dynamics of them by using simple physical system. Then we try to apply the optimization techniques, in order to tune parameters of those models of that skill.
4.1 Modelling and Controlling the Dynamics of the Low-Level Skills
Many popular approaches used for controlling the balance of bipedal locomotion are based on the Zero Momentum Point (ZMP) stability indicator and inverted pendulum model. ZMP cannot generate reference walking trajectories directly but it can indicate whether generated walking trajectories will keep the balance of a robot or not. Kajita et al. assumed that biped walking is a problem of balancing a cart-table model [8], since in the single supported phase, human walking can be represented as the Cart-table model.
Biped walking can be modeled through the movement of ZMP and CoM. The robot is in balance when the position of the ZMP is inside the support polygon. When the ZMP reaches the edge of this polygon, the robot loses its balance. Cart-table model has some assumptions and simplifications in its model. One the major drawback of the cart-table model is its consideration the height of the robot fixed during it movement, which is not true for many soccer low-level skills such as running or kicking, Therefore we used the inverted pendulum model which does not have this issue. Figure 2 shows how robot dynamics is modeled by an inverted pendulum and its schematic view.
4.2 Omni-Directional Biped Locomotion
This section briefly presented the design of the locomotion controllers to enable the robot with an omni-directional walking. We use this design to implement walking approached by using both inverted pendulum and cart –table model. The extended details of our approach can be found in [13].
Developing an omni-directional biped locomotion is a complex task made of several components. In order to get a functional omnidirectional walk, it is necessary to decompose it in several modules and address each module independently. Each modules is explains in the following.
- ZMP Trajectory Generator - In this module the ZMP generated using the desired velocities. This computation takes into account only the linear component of the walk, which means to walk in any direction always looking to the same direction, like diagonal walk.
- Foot Planner One of the drawbacks of the linear inverted pendulum model is the need for a constant height. We can improve this by adjusting the CoM height using the length of the leg, support foot position and the ground projection of the CoM. In [12] a detailed explanation is given.
- CoM Trajectory generator This module is responsible to Generate the CoM by using the dynamics equations of inverted pendulum model [14], or cart-table model [12] [13].
- Swing Trajectory Generator This module is responsible to generate a trajectory for the swing foot. It uses the cycloid parametric equation to generate the desired trajectory.
- Feet Frame Computation After computing the support foot (ZMP) position, CoM position and swing trajectory feet position has to be computed taking into account if it is in double support phase or single (left or right) support phase. This module is responsible for computing the position and orientation of both feet relative to the CoM frame.
Active Balance - This module is where the balance of the humanoid during locomotion is controlled in order to maintain it stable. The inverted pendulum model has some simplifications in biped walking dynamics modeling; in addition, there is inherent noise in leg's actuators. Therefore, keeping walk balance generated by inverted pendulum model cannot be guaranteed. In order to reduce the risk of falling during walking, an active balance technique is applied. The detailed explanation of this module can be found in [12].
The Active balance module tries to keep an upright trunk position by decreasing variation of trunk angles. One PD controller is designed to control the trunk angle to be the desired trunk pitch angle. An inertial measurement unit which is included in the robot body gives the trunk inclination angle. When the trunk angle is not the desired pitch angle, instead of considering a coordinate frame attached to the trunk of the biped robot, position and orientation of the feet are calculated with respect to a coordinate frame, which is attached to the CoM position and the Z axes always has the predefined pitch angle to the ground plane. For example if the trunk pitch offset is assumed to be zero, the Z axes keeps always perpendicular to the ground plane.
The PD controller calculates the rotation angle based on the difference between the current trunk inclination and the desired trunk pitch angle. The calculated rotation angle is a portion of this difference and the coordinate frame rotates with the calculated rotation angle. By using this transformation, the controller tries to keep the Z axis of the coordinate frame in a desired angle to the ground plane. The foot position is calculated by using the rotated coordinated frame, the feet orientation also tries to be kept parallel to the ground.
The Transformation formulation is presented in equation (4).
$$Foot = T_{Foot}^{CoM}(pitchAng, rollAng) \times Foot$$ (4)
The pitchAng and rollAng are assumed to be the angles calculated by the PID controller around y and x axis respectively. Figure 3 shows the architecture of the active balance unit when the trunk pitch offset is assumed to be zero.
4.3 Omni-Directional Running
Many researchers, up to now, have modeled the biped walking by considering the height of Center of Mass (CoM) as a fixed constant, other biomechanical studies show that the CoM height is variant during walking and running [17]. The shape of CoM height trajectory is important for energy consumption, and it varies differently for various speeds and step length ranges. Recently, we have showed this fact in our work [16]. Although cart-table model is widely used in robotics, but robots often need to keep their knees bent in order to keep the height of CoM fixed, as the constraint of this model. Therefore, the change of the CoM height is important both for walking and running. Figure 4 and figure 5 shows a planar view of the human walking and a walking generated by a cart-table model, respectively.
4.4 Humanoid Kick with Controlled Distance
We investigate the learning of a flexible humanoid robot kick controller, i.e., the controller should be applicable for multiple contexts, such as different kick distances, initial robot position with respect to the ball or both. Current approaches typically tune or optimise the parameters of the biped kick controller for a single context, such as a kick with longest distance or a kick with a specific distance. Hence our research question is "how can we obtain a flexible kick controller that controls the robot (near) optimally for a continuous range of kick distances?". The goal is to find a parametric function that given a desired kick distance, outputs the (near) optimal controller parameters. We achieve the desired flexibility of the controller by applying a contextual policy search method. With such a contextual policy search algorithm, we can generalize the robot kick controller for different distances, where the desired distance is described by a real-valued vector.
Figure 8 shows an example of an initial and final stance for the kick behavior. Our movement pipeline is composed of two main parts: a kick controller, which receives parameters and converts them into joint commands for the robot's servos; and a policy function, which maps a given context s for a specific kick distance into the corresponding parameter vector . The pipeline for the kick task, whose context is the kick distance s with a straight kick direction with respect to the torso, is shown in Figure 9.
5. High-Level Decisions and Coordination
Flexible Tactics has always been one of the major assets of FC Portugal teams. FC Portugal 3D is capable of using several different formations and for each formation players may be instantiated with different player types. The management of formations and player types is based on SBSP – Situation Based Strategic Positioning algorithm [1, 4]. Player's abandon their strategic positioning when they enter a critical behavior: Ball Possession or Ball Recovery. This enables the team to move in a quite smooth manner, keeping the field completely covered.
The high-level decision uses the infrastructure presented in the section 3. Several new types of actions are currently being considered taking in consideration the new opportunities opened by the 3D environment of the new simulator. We also have adapted our previous researched methodologies to the new 3D environment:
Strategy for a Competition with a Team with Opposite Goals [1, 4, 5, 21];
Concepts of Tactics, Formations and Player Types [1, 3, 4, 21];
Distinction between Active and Strategic Situations [1, 4];
Situation Based Strategic Positioning (SBSP) [1, 4, 5];
Dynamic Positioning and Role Exchange (DPRE) [1, 4, 5];
Visual Debugging and Analysis Tools [1, 3, 22];
Optimization based Low-Level Skills [1, 3, 26, 27].
Standard Language to Coach a (Robo)Soccer Team[2,3];
Intelligent Communication using a Communicated World State [1, 3, 5];
Flexible Set-plays for coordinating robosoccer teams [23].
In 2012, 2013 and 2014, our research was mostly concerned in developing optimization based low level skills for the humanoid agent and robust mid-level skills. The high-level layers of the team for 2016 will be adapted to be used in the humanoid simulator (these methodologies have already been adapted to our Simulation 2D, Simulation 3D with spheres model, small-size, middle-size [24] and rescue teams [25]).
6. Learning Framework
For developing a learning framework, it is way better to run the simulation as fast as the CPU can. By using Syncmode, the simspark simulator only waits for the agent commands and a synchronize message which signals the end of the agent cycle. After it receives all the agents Sync message, the server processes all the commands and proceeds to the next cycle. In addition to simulation speed time improvement, it can also be used to detect strange cycle times from the agents. We have developed our agent to use Sync Mode to improve speed of the optimization process.
The process of optimization is, the 3D soccer server, and a Matlab program. All of our optimization algorithm such as CMA-ES, GA and etc. are developed using Matlab code because of the facility that Matlab prepares for mathematical programming. After connection, the optimization agent sends the required optimization data to the Matlab program, that, afterwards, uses the agent as a server to compute the cost function of the individuals. By its turn, the agent, in order to compute the value of the cost function, runs a simulation in the soccer server, and returns the computed value to the Matlab program. These interactions last until the optimization process ends. Figure 11 illustrate the interaction among the different applications.
Conclusions
Robust low-level skills have been developed for the NAO humanoid model, The results of low-level skills have already tested and validated on the real NAO robot, since it is based on the physical modeling of the dynamics of biped locomotion it is very robust and with minimal adaptation was used on the NAO robot. Using optimization and learning techniques, enabling us to continue the research in strategical reasoning and coordination methodologies that should be the focus of the simulation leagues inside RoboCup. Also the extended flexibility of omnidirectional kicks and walks will enable a more cooperative game style.
Future work will be concerned in extending the optimization methodology for skills sequences and on developing coordination methodologies enabling teams of humanoid robots to play robosoccer games in a robust and flexible manner.
Almost all of our research on high-level flexible coordination methodologies is directly applicable to the 3D league and the increase in the number of elements of the each team is very welcome, enabling coordination methodologies to be useful in this league. FC Portugal started its participation in the SPL - Standard Platform League in 2011. The SPL code of the team is entirely made from scratch based on the Simulation 3D code. Thus, future work will be on bridging the gap between simulation and robotics by developing a more realistic NAO model in Simspark enabling better portability of the simulated code to the real robot.
Acknowledgements
This work was partially supported by the Portuguese National Foundation for Science and Technology: SFRH/BD/66597/2009 and SFRH/BD/81155/2011.
References
- Luis Paulo Reis and Nuno Lau, FC Portugal Team Description: RoboCup 2000 Simulation League Champion, In P. Stone, T. Balch and G. Kraetzschmar editors, RoboCup-2000: Robot Soccer World Cup IV, LNAI 2019, pp 29-40, Springer, 2001.
- Luís Paulo Reis and Nuno Lau, COACH UNILANG A Standard Language for Coaching a (Robo) Soccer Team, in Andreas Birk, Silvia Coradeschi and Satoshi Tadokoro, editors, RoboCup2001 Symposium: Robot Soccer World Cup V, Springer Verlag LNAI, Vol. 2377, pp. 183-192, Berlin, 2002.
- Nuno Lau and Luis Paulo Reis, FC Portugal Homepage, [online] available at: Http://www.ieeta.pt/robocup, consulted on February 2010.
- Luis Paulo Reis, Nuno Lau and Eugénio C. Oliveira, Situation Based Strategic Positioning for Coordinating a Team of Homogeneous Agents in M.Hannebauer, et al.eds, Balancing Reactivity and Social Deliberation in MAS – From RoboCup to Real-World Applications, Springer LNAI, Vol. 2103, pp. 175-197, 2001
- Nuno Lau and Luis Paulo Reis, FC Portugal High-level Coordination Methodologies in Soccer Robotics, Robotic Soccer, Book edited by Pedro Lima, Itech Education and Publishing, Vienna, Austria, pp. 167-192, December 2007, ISBN 978-3-902613-21-9
- Hugo Picado, Marcos Gestal, Nuno Lau, Luís Paulo Reis, Ana Maria Tomé, Automatic Generation of Biped Walk Behavior Using Genetic Algorithms. In J.Cabestany et al. (Eds.), 10th International Work-Conference on Artificial Neural Networks, IWANN 2009, Springer, LNCS Vol. 5517, pp. 805-812, Salamanca, Spain, June 10-12, 2009
- Hugo Marques, Nuno Lau and Luís Paulo Reis, FC Portugal 3D Simulation Team: Architecture, Low-Level Skills and Team Behaviour Optimized for the New RoboCup 3D Simulator, Reis, L.P. et al. eds, Proc. Scientific Meeting of the Portuguese Robotics Open 2004, FEUP Editions, C.Colectâneas, Vol. 14, pp.31-37, April, 23-24, 2004
- S. Kajita, F. Kanehiro, K. Kaneko, K. Yokoi, and H. Hirukawa, The 3D linear inverted pendulum mode: a simple modeling for a biped walking pattern generation, in IEEE/RSJ International Conference on Intelligent Robots and Systems, 2001, pp. 239–246.
- M. Vukobratovic, D. Stokic, B. Borovac, and D. Surla. Biped Locomotion: Dynamics, Stability, Control and Application. Springer Verlag, 1990, p. 349.
- Rui Ferreira, Luís Paulo Reis, António Paulo Moreira, Nuno Lau, Development of an Omnidirectional Kick for a NAO Humanoid RobotIn Advances in Artificial Intelligence – IBERAMIA 2012, Lecture Notes in Computer Science, Volume 7637, Springer, 2012, pp.571-580
- E. Domingues, N. Lau, B. Pimentel, N. Shafii, L. P. Reis, A. J. R. Neves, Humanoid Behaviors: From Simulation to a Real Robot, 15th Port. Conf. Artificial Intelligence, EPIA 2011, LNAI 7367, pp 352-364, 2011.
- Shafii, N., Abdolmaleki, A., Ferreira, R., Lau, N., & Reis, L. P. Omnidirectional Walking and Active Balance for Soccer Humanoid Robot. In Progress in Artificial Intelligence (pp. 283-294), 2013.
- Ferreira, Rui, Nima Shafii, Nuno Lau, Luis Paulo Reis, and Abbas Abdolmaleki. Diagonal walk reference generator based on Fourier approximation of ZMP trajectory. In Autonomous Robot Systems (Robotica), 2013 13th International Conference on, pp. 1-6, 2013.
- Nima Shafii, Nuno Lau, Luis Paulo Reis. Learning to Walk Fast: Optimized Hip Height Movement for Simulated and Real Humanoid Robots, Journal of Intelligent & Robotic Systems, Springer, 2015.
- Nima Shafii, Nuno Lau, Luis Paulo Reis. Learning a fast walk based on ZMP control and hip height movement. In 2014 IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC), pp. 181-186, 2014.
- Nima Shafii, Nuno Lau, Luis Paulo Reis. Generalized Learning to Create an Energy Efficient ZMP-Based Walking. RoboCup-2014: Robot Soccer World Cup XVIII, Lecture Notes in Computer Science, Springer, 2015.
- Kuo, A. D., Donelan, J. M., & Ruina, A. Energetic consequences of walking like an inverted pendulum: step-to-step transitions. Exercise and Sport Sciences Reviews, 33(2), 88–97, 2007.
- Shafii, N., Reis, L. P., & Lau, N.. Biped Walking using Coronal and Sagittal Movements based on Truncated Fourier Series. In Javier Ruiz-del-Solar, E. Chown, & Paul-Gerhard (Eds.), RoboCup 2010: Robot Soccer World Cup XIV , pp. 324–335, 2011.
- N. Shafii, M. H. Javadi, B. Kimiaghalam. A Truncated Fourier Series with Genetic Algorithm for the control of Biped Locomotion. In: Proceeding of the 2009 IEEE/ASME International Conference on advanced intelligent Mechatronics, pp. 1781--1785, 2009.
- Nima Shafii, Nuno Lau, Luis Paulo Reis. Generalized Optimization and Reinforcement Learning for Creating an Energy Efficient ZMP Walker, In Proceedings of the 2014 Dynamic Walking, ETH Zurich, 2014.
- JoãoCerto, Nuno Lau e Luís Paulo Reis, A Generic Multi-Robot Coordination Strategic Layer, RoboComm 2007 - First International Conference on Robot Communication and Coordination, Athens, Greece, October 15-17, 2007.
- Nuno Lau, Luís Paulo Reis e JoãoCerto, Understanding Dynamic Agent's Reasoning, In Progress in Artificial Intelligence, 13th Port. Conf. on AI, EPIA 2007, Guimarães, Portugal, December 3-6, 2007, Springer LCNS, Vol. 4874, pp. 542-551, 2007
- Luís Mota e Luís Paulo Reis, Setplays: Achieving Coordination by the appropriate Use of arbitrary Pre-defined Flexible Plans and inter-robot Communication, RoboComm 2007 - First International Conf. on Robot Communication and Coordination, Athens, Greece, October 15-17, 2007
- Nuno Lau, LuísSeabra Lopes, G. Corrente and Nelson Filipe, Multi-Robot Team Coordination Through Roles, Positioning and Coordinated Procedures, Proc.IEEE/RSJ Int. Conf. on Intelligent Robots and Systems – IROS 2009, St. Louis, USA, Oct. 2009
- Luís Paulo Reis, Nuno Lau, Francisco Reinaldo, Nuno Cordeiro and João Certo. FC Portugal: Development and Evaluation of a New RoboCup Rescue Team. 1st IFAC Workshop on Multivehicle Systems (MVS'06), Salvador, Brazil, October 2 – 3, 2006
- Luis Rei, Luis Paulo Reis and Nuno Lau, Optimizing a Humanoid Robot Skill, Robótica 2011 11th International Conference on Mobile Robots and Competitions, pp. 78-83, Lisbon, Portugal, 2011 (Best Paper Award)
- Luis Cruz, Luis Paulo Reis, Nuno Lau, Armando Sousa, Optimization Approach for the Development of Humanoid Robots' Behaviors, In Advances in Artificial Intelligence – IBERAMIA 2012, Lecture Notes in Computer Science, Volume 7637, Springer, 2012, pp. 491-500.
- Abbas Abdolmaleki, David Simões, Nuno Lau, Luis Paulo Reis. Learning a Humanoid Kick With Controlled Distance. RoboCup 2016: Robot World Cup XX, July 2016
- Abbas Abdolmaleki, Nuno Lau, Luis Paulo Reis, Gerhard Neumann. Non-parametric contextual stochastic search. (2016) IEEE International Conference on Intelligent Robots and Systems, South Korea, p. 2643-2648, November 2016