Scinovex
article Open AccessTop 1% cited

Temporal difference learning and TD-Gammon

Communications of the ACM · 1995 · Vol. 38(3) · pp. 58–68
Gerald Tesauro

Abstract

Ever since the days of Shannon's proposal for a chess-playing algorithm [12] and Samuel's checkers-learning program [10] the domain of complex board games such as Go, chess, checkers, Othello, and backgammon has been widely regarded as an ideal testing ground for exploring a variety of concepts and approaches in artificial intelligence and machine learning. Such board games offer the challenge of tremendous complexity and sophistication required to play at expert level. At the same time, the problem inputs and performance measures are clear-cut and well defined, and the game environment is readily automated in that it is easy to simulate the board, the rules of legal play, and the rules regarding when the game is over and determining the outcome.

Artificial Intelligence in GamesReinforcement Learning in RoboticsAdvanced Bandit Algorithms ResearchSophisticationComputer scienceVariety (cybernetics)Artificial intelligenceIdeal (ethics)Domain (mathematical analysis)Outcome (game theory)Machine learningMathematicsMathematical economics
Citations
1,472
FWCI
25.28
field-weighted impact
References
7
Percentile
100%
vs. same field & year
Citations per year
References
Learning to Predict by the Methods of Temporal Differences
Machine Learning · 1988 · 3,908 citations
Practical issues in temporal difference learning
Machine Learning · 1992 · 795 citations
Multilayer feedforward networks are universal approximators
Neural Networks · 1989 · 20,841 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.