Practical Reinforcement Learning -Experiences in Lot Scheduling Application
Author(s) -
Hannu Rummukainen,
Jukka K. Nurminen
Publication year - 2019
Publication title -
ifac-papersonline
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.308
H-Index - 72
eISSN - 2405-8971
pISSN - 2405-8963
DOI - 10.1016/j.ifacol.2019.11.397
Subject(s) - reinforcement learning , computer science , scheduling (production processes) , a priori and a posteriori , benchmarking , mathematical optimization , stochastic control , parameterized complexity , artificial intelligence , job shop scheduling , artificial neural network , q learning , optimal control , machine learning , mathematics , schedule , algorithm , economics , operating system , epistemology , philosophy , management
With recent advances in deep reinforcement learning, it is time to take another look at reinforcement learning as an approach for discrete production control. We applied proximal policy optimization (PPO), a recently developed algorithm for deep reinforcement learning, to the stochastic economic lot scheduling problem. The problem involves scheduling manufacturing decisions on a single machine under stochastic demand, and despite its simplicity remains computationally challenging. We implemented two parameterized models for the control policy and value approximation, a linear model and a neural network, and used a modified PPO algorithm to seek the optimal parameter values. Benchmarking against the best known control policy for the test case, in which Paternina-Arboleda and Das (2005) combined a base-stock policy and an older reinforcement learning algorithm, we improved the average cost rate by 2 %. Our approach is more general, as we do not require a priori policy parameters such as base-stock levels, and the entire policy is learned.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom