Long short-term memory (LSTM) is a type of recurrent neural network that processes sequential data using an internal memory state and learned gates. Within machine learning, it provides a way to retain information across many processing steps rather than relying only on recent inputs. Its central design addresses the vanishing gradient problem, which makes long-range dependencies difficult to learn in conventional recurrent networks. Sepp Hochreiter and Jürgen Schmidhuber introduced the architecture in their 1997 paper “Long Short-Term Memory.” (doi.org)
Origins and motivation
A recurrent network repeatedly updates its state as it receives successive inputs. Training requires determining how earlier computations contributed to later errors. Repeated multiplication of derivatives can make the resulting gradients extremely small or large, producing vanishing gradients or the exploding gradient problem. Consequently, information relevant to a prediction may become difficult to learn when separated from that prediction by many intervening steps. (arxiv.org)
The original LSTM introduced memory cells with a specially structured recurrent connection and multiplicative input and output gates. Its “constant error carousel” provided a route through which error signals could persist across time. Experiments demonstrated learning across delays exceeding 1,000 steps on selected artificial tasks; this was a benchmark result, not a guaranteed memory length for every application. (doi.org)
Felix Gers, Schmidhuber, and Fred Cummins subsequently introduced an adaptive forget gate, described in a 2000 journal article. It enabled cells to discard stored information when processing continuous streams without explicitly marked reset points. The familiar three-gate LSTM therefore differs from the original 1997 formulation. (pubmed.ncbi.nlm.nih.gov)
Cell structure and equations
A standard LSTM maintains two vectors: a cell state, , which carries stored information, and a hidden state, , which exposes a gated representation to subsequent computations. At step , both are updated using the current input and previous states. The gates use the logistic sigmoid as an activation function, producing continuous values between zero and one rather than binary switches. (docs.pytorch.org)
A common formulation without peephole connections is:
Here, and are learned weight matrices, denotes biases, and means elementwise multiplication. The input gate controls writing, the forget gate controls retention, and the output gate controls exposure of the cell state. The candidate vector supplies new content. (docs.pytorch.org)
The additive cell update is crucial. Along the direct state-to-state path, retained information is multiplied by the forget gate rather than repeatedly transformed through a full nonlinear recurrent mapping. Gates near one can preserve that path over many steps. This improves gradient propagation but does not guarantee unlimited retention or eliminate every optimization difficulty. (doi.org)
Training and sequence processing
LSTMs are commonly trained using backpropagation through time, which applies backpropagation to the network unfolded across a sequence. A task-specific loss function measures prediction error, and gradient descent adjusts the shared parameters. For next-step prediction, training can maximize the likelihood of observed sequences by minimizing their negative log likelihood. (arxiv.org)
Although the memory pathway helps with vanishing gradients, excessively large derivatives can still occur. Gradient clipping limits their magnitude during optimization; both general recurrent-network research and LSTM sequence-generation experiments document its use. Clipping addresses numerical instability rather than expanding the cell’s representational capacity. (arxiv.org)
The same transition parameters are reused at each step, allowing an LSTM to process sequences of varying lengths. Its hidden representations can feed an output layer at every position or provide an encoded representation for another network. Multilayer arrangements stack recurrent layers, forming a deep learning model whose higher layers receive representations from lower layers. (arxiv.org)
Variants and applications
A bidirectional LSTM processes a sequence in both directions and combines their representations. It therefore incorporates later as well as earlier context, unlike a purely forward recurrent model. Projected LSTMs apply a learned projection to the hidden output, allowing its dimension to differ from the cell-state dimension. Peephole variants additionally connect cell states directly to gates. (docs.pytorch.org)
LSTMs have been applied to speech recognition and handwriting recognition. In natural language processing, their uses include language modeling and text generation. Alex Graves demonstrated generation of text and online handwriting by predicting successive elements, including handwriting conditioned on supplied text. (arxiv.org)
A prominent sequence-to-sequence approach used one multilayer LSTM to encode an input and another to decode an output. Ilya Sutskever, Oriol Vinyals, and Quoc Le demonstrated this architecture for English-to-French machine translation in 2014, establishing a general method for mapping between variable-length sequences. (arxiv.org)
Related architectures and computational limits
The gated recurrent unit (GRU) is a related gated architecture that combines memory and hidden-state roles rather than maintaining the LSTM’s separate cell state. A 2014 comparison found GRUs comparable to LSTMs on the evaluated speech and music tasks, without establishing a universally superior unit. (arxiv.org)
LSTM computation remains sequential: each recurrent state depends on its predecessor. This constrains parallel computation across sequence positions. The Transformer architecture, introduced in 2017, instead uses an attention mechanism without recurrence in its original formulation, enabling greater training parallelism. The architectures consequently differ both in how they represent sequence context and in their computational dependencies. (arxiv.org)