Temporal Difference Learning and TD-Gammon
Author(s) -
G. Tesauro
Publication year - 1995
Publication title -
icga journal
Language(s) - English
Resource type - Journals
eISSN - 2468-2438
pISSN - 1389-6911
DOI - 10.3233/icg-1995-18207
Subject(s) - computer science , astrophysics , physics
We provide an abstract, selectively u§ing the author's formulations: "The article presents a game-learning program called TD-GAMMON. TD-GAMMON is a neural network that trains itself to be an evaluation function for the game of backgammon by playing against itself and learning from the outcome. It was not developed to surpass all previous computer programs in backgammon; rather, its purpose was to explore some new ideas and approaches to traditional problems in reinforcement learning.
Accelerating Research
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom
Address
John Eccles HouseRobert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom