Brainstormers 2004 - Team Description

M. Riedmiller, A. Merke, D. Withopf

AG Neuroinformatik, Universität Osnabrück, 49069 Osnabrück, Germany


Abstract The main interest behind the Brainstormers' effort in the robocup soccer domain is to develop and to apply machine learning techniques in complex domains. Especially, we are interested in reinforcement learning methods, where the training signal is only given in terms of success or failure. Our final goal is a learning system, where we only plug in 'win the match' - and our agents learn to generate the appropriate behaviour. Unfortunately, even from very optimistic complexity estimations it becomes obvious, that in the soccer domain, both conventional solution techniques and also advanced today's reinforcement learning techniques come to their limit - there are more than $(10^8 \times 50)^{23}$ different states and more than $(1000)^{300}$ different policies per agent per half time. This paper describes the new architecture of the Brainstormers team, the improved self-localization using particle filters and the extensions of the learning algorithm to simultaniously learn with and without ball behaviors.

1 Design Principles

Both our 2D and our 3D teams are based on the following principles:

  • two main modules: world module and decision making module
  • input to the decision module is the approximate, complete world state
  • the soccer environment is modelled as an Markovian Decision Process (MDP)
  • decision making is organized in complex and less complex behaviours
  • a steadily growing part of the behaviours is learned by Reinforcement Learning methods (2004: intercept, go2pos, shoot, 1-versus-1, attack play: running free, dribbling, passing, scoring)
  • modern AI methods are applied wherever possible and useful (e.g particle filters are used for improved self localisation)

2 Reinforcement Learning of Team Strategies

Many of the basic skills of our team like intercept-ball, goto-position and kick have been implemented using Reinforcement Learning approaches. These skills can be described as single-agent markov decision processes (MDP). For more information on learning single-agent skills see [4].

Our main research interest currently lies in the field of multi-agent markov decision processes (MMDP). Learning a team strategy can be described as such a MMDP with individual learners. As already stressed in [5] there is no guarantee that an optimal strategy is learned in such a case. In this section we describe the learning algorithm that was used to learn the positioning of players without ball and also the actions of the player with ball.

Fig. 1. The new behavior architecture
Fig. 1. The new behavior architecture

References

  1. Betke, M., Gurvits, L.: Mobile robot localization using landmarks. IEEE Transactions on Robotics and Automation 13 (1997) 251–263
  2. Fox, D., Burgard, W., Dellaert, F., Thrun, S.: Monte carlo localization: Efficient position estimation for mobile robots. In: Proceedings of the Sixteenth National Conference on Artificial Intelligence (AAAI'99). (1999)
  3. Fox, D., Thrun, S., Burgard, W., Dellaert, F.: Particle filters for mobile robot localization. In Doucet, A., de Freitas, N., Gordon, N., eds.: Sequential Monte Carlo Methods in Practice, New York, Springer (2001)
  4. Riedmiller, M., Merke, A., Meier, D., Hoffmann, A., Sinner, A., Thate, O., Kill, C., Ehrmann, R.: Karlsruhe brainstormers - a reinforcement learning way to robotic soccer. In Jennings, A., Stone, P., eds.: RoboCup-2000: Robot Soccer World Cup IV, LNCS. Springer (2000)
  5. Merke, A., Riedmiller, M.: Karlsruhe brainstormers a reinforcement learning way to robotic soccer ii. In Birk, A., Coradeschi, S., Tadokoro, S., eds.: RoboCup-2001: Robot Soccer World Cup V, LNCS. Springer (2001) 322–327
  6. Riedmiller, M., Braun, H.: A direct adaptive method for faster backpropagation learning: The RPROP algorithm. In Ruspini, H., ed.: Proceedings of the IEEE International Conference on Neural Networks (ICNN), San Francisco (1993) 586 – 591
  7. Lauer, M., Riedmiller, M.: An algorithm for distributed reinforcement learning in cooperative multi-agent systems. In: Proceedings of International Conference on Machine Learning, ICML '00, Stanford, CA (2000) 535–542