Cross Entropy Optimization of Action Modification Policies for Continuous-Valued MDPs
Author(s) -
Kamelia Mirkamali,
Lucian Buşoniu
Publication year - 2020
Publication title -
ifac-papersonline
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.308
H-Index - 72
eISSN - 2405-8971
pISSN - 2405-8963
DOI - 10.1016/j.ifacol.2020.12.2292
Subject(s) - markov decision process , mathematical optimization , bellman equation , entropy (arrow of time) , computer science , basis (linear algebra) , novelty , action (physics) , basis function , integrator , principle of maximum entropy , function (biology) , markov process , mathematics , artificial intelligence , bandwidth (computing) , evolutionary biology , mathematical analysis , theology , computer network , biology , statistics , physics , quantum mechanics , geometry , philosophy
We propose an algorithm to search for parametrized policies in continuous state and action Markov Decision Processes (MDPs). The policies are represented via a number of basis functions, and the main novelty is that each basis function corresponds to a small, discrete modification of the continuous action. In each state, the policy chooses a discrete action modification associated with a basis function having the maximum value at the current state. Empirical returns from a representative set of initial states are estimated in simulations to evaluate the policies. Instead of using slow gradient-based algorithms, we apply cross entropy method for updating the parameters. The proposed algorithm is applied to a double integrator and an inverted pendulum problem, with encouraging results.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom