z-logo
open-access-imgOpen Access
Co-evolution of Shaping Rewards and Meta-Parameters in Reinforcement Learning
Author(s) -
Stefan Elfwing,
Eiji Uchibe,
Kenji Doya,
Henrik I. Christensen
Publication year - 2008
Publication title -
adaptive behavior
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.286
H-Index - 54
eISSN - 1741-2633
pISSN - 1059-7123
DOI - 10.1177/1059712308092835
Subject(s) - reinforcement learning , action selection , softmax function , task (project management) , meta learning (computer science) , computer science , artificial intelligence , foraging , animal learning , selection (genetic algorithm) , reinforcement , action (physics) , machine learning , cognitive psychology , artificial neural network , psychology , engineering , social psychology , ecology , systems engineering , neuroscience , perception , biology , physics , quantum mechanics
Digital Object Identifier: 10.1177/1059712308092835In this article, we explore an evolutionary approach to the optimization of potential-based shaping rewards and meta-parameters in reinforcement learning. Shaping rewards is a frequently used approach to increase the learning performance of reinforcement learning, with regards to both initial performance and convergence speed. Shaping rewards provide additional knowledge to the agent in the form of richer reward signals, which guide learning to high-rewarding states. Reinforcement learning depends critically on a few meta-parameters that modulate the learning updates or the exploration of the environment, such as the learning rate α, the discount factor of future rewards γ, and the temperature τ that controls the trade-off between exploration and exploitation in softmax action selection. We validate the proposed approach in simulation using the mountain-car task. We also transfer shaping rewards and meta-parameters, evolutionarily obtained in simulation, to hardware, using a robotic foraging task

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom