Scinovex
articleTop 10% cited

Learning to Forget: Continual Prediction with LSTM

Neural Computation · 2000 · Vol. 12(10) · pp. 2451–2471
Felix A. GersJürgen SchmidhuberFred Cummins

Abstract

Long short-term memory (LSTM; Hochreiter & Schmidhuber, 1997) can solve numerous tasks not solvable by previous learning algorithms for recurrent neural networks (RNNs). We identify a weakness of LSTM networks processing continual input streams that are not a priori segmented into subsequences with explicitly marked ends at which the network's internal state could be reset. Without resets, the state may grow indefinitely and eventually cause the network to break down. Our remedy is a novel, adaptive "forget gate" that enables an LSTM cell to learn to reset itself at appropriate times, thus releasing internal resources. We review illustrative benchmark problems on which standard LSTM outperforms other RNN algorithms. All algorithms (including LSTM) fail to solve continual versions of these problems. LSTM with forget gates, however, easily solves them, and in an elegant way.

Domain Adaptation and Few-Shot LearningNeural Networks and ApplicationsAdversarial Robustness in Machine LearningBenchmark (surveying)Reset (finance)Computer scienceRecurrent neural networkArtificial intelligenceA priori and a posterioriState (computer science)Artificial neural networkMachine learningAlgorithm

MeSH terms

AlgorithmsMemory, Short-TermNeural Networks, ComputerNonlinear Dynamics

Funding

  • Schweizerischer Nationalfonds zur Förderung der Wissenschaftlichen Forschung
Citations
5,306
FWCI
10.66
field-weighted impact
References
21
Percentile
98%
vs. same field & year
Citations per year
References
Long Short-Term Memory
Neural Computation · 1997 · 95,078 citations
Learning long-term dependencies in NARX recurrent neural networks
IEEE Transactions on Neural Networks · 1996 · 782 citations
Learning long-term dependencies with gradient descent is difficult
IEEE Transactions on Neural Networks · 1994 · 8,303 citations
LSTM recurrent networks learn simple context-free and context-sensitive languages
IEEE Transactions on Neural Networks · 2001 · 724 citations
Citation Network

How this paper connects to the literature. Drag to explore, click any node to open that paper.