A recurrent neural network (RNN) is a family of artificial neural networks designed to process ordered sequences. Its defining feature is recurrence: information from one processing step influences subsequent steps through an internal state. The same learned update rule is applied repeatedly, allowing the network to handle variable-length inputs without allocating separate parameters to every position. RNNs are used in machine learning for tasks involving text, speech, and time series. A sequence position need not represent physical time; it may instead correspond to a word, character, or other ordered element. (deeplearningbook.org)
Structure and computation
In a simple RNN, the input vector and previous hidden state determine the new hidden state:
Here, each is a learned weight matrix, is a bias vector, and is an activation function, commonly the hyperbolic tangent. An initial state supplies the starting condition, often using zeros. A separate output layer can transform the hidden state into a prediction. Because the recurrent parameters are reused at every position, increasing sequence length increases computation but not the number of parameters in the recurrent layer. (docs.pytorch.org)
The hidden state is a learned, generally lossy representation of the preceding inputs, not an exact record of them. Which information it retains depends on the training objective. “Unrolling” the recurrence represents successive applications as an expanded computational graph, with shared parameters connecting the steps. This distinguishes recurrence from simply presenting a fixed window of observations to a feedforward network. (deeplearningbook.org)
Recurrent layers can also be stacked. In a stacked network, one layer’s sequence of hidden states becomes the next layer’s input, adding representational depth alongside the repeated computation across sequence positions. (docs.pytorch.org)
Training and gradient difficulties
RNN training commonly uses backpropagation through time (BPTT), which applies backpropagation to the unrolled network. A loss function measures prediction errors at selected positions or over the complete sequence. Contributions from all uses of a shared parameter are accumulated, and an optimization method such as stochastic gradient descent updates the parameters. Truncated BPTT restricts gradient propagation to a limited span, reducing computational and memory requirements while limiting how far learning signals travel. (deeplearningbook.org)
Long sequences create important optimization difficulties. Through repeated multiplication by Jacobian matrices, gradients may become extremely small or large. The vanishing gradient problem makes learning relationships between widely separated inputs difficult; the exploding gradient problem can produce unstable parameter updates. Gradient clipping limits gradient magnitude before an update and addresses exploding gradients, but does not by itself resolve vanishing gradients. Gated architectures alter the recurrent computation to provide more effective paths for information and error signals. (arxiv.org)
Gated recurrent architectures
Long short-term memory (LSTM), introduced by Sepp Hochreiter and Jürgen Schmidhuber in 1997, was developed to address difficulties in learning long-term dependencies. It introduces memory cells and multiplicative gates that regulate information flow. Common later formulations distinguish a cell state from the exposed hidden state and use input, forget, and output gates to control writing, retention, and reading. The additive cell-state pathway can preserve information and facilitate gradient propagation across many steps, although it does not guarantee successful learning of every long-range dependency. (bioinf.jku.at)
The gated recurrent unit (GRU), proposed in 2014, uses update and reset gates. Its update gate controls the balance between retaining the previous state and incorporating a candidate state; its reset gate regulates how previous information contributes to that candidate. Unlike a standard LSTM, it does not maintain a separate cell state. These differences generally give a GRU fewer parameters than an LSTM with the same input and hidden dimensions, but do not establish a universally superior architecture. (arxiv.org)
Sequence configurations and applications
An RNN can return an output at every position or a single representation after processing a sequence. These configurations support sequence labeling and sequence classification, respectively. Libraries also distinguish the recurrent cell—the computation for one step—from the layer that applies it across a sequence. States may be passed between successive input segments, enabling continued processing without restarting the recurrence. (tensorflow.org)
A bidirectional RNN processes a sequence in both directions and combines their representations. Each position can therefore incorporate preceding and following context. This requires access to the relevant future inputs, unlike a forward-only network operating causally on an incoming stream. (deeplearningbook.org)
In natural language processing, recurrent language models predict subsequent symbols from earlier context. RNNs have also been applied to speech recognition and handwriting processing. An encoder–decoder architecture uses one network to represent an input sequence and another to generate an output sequence, potentially of a different length. The 2014 RNN encoder–decoder study demonstrated this approach for phrase representations used in machine translation. (deeplearningbook.org)
Historical development and computational trade-offs
Jeffrey Elman’s 1990 simple recurrent network became an influential model for learning temporal structure. Later developments included LSTM and GRU architectures, followed by combinations of recurrent encoder–decoders with an attention mechanism. Attention allows a decoder to access different input representations rather than relying exclusively on a single fixed-length summary. (crl.ucsd.edu)
The Transformer architecture, introduced in 2017, dispensed with recurrence in favor of attention-based sequence processing. Conventional RNN states must be computed sequentially, restricting parallel computation across positions. Transformers permit greater parallelism during training and provide shorter computational paths between distant positions. Conversely, a forward RNN carries a fixed-size state as it processes new observations; this compact representation also limits how much earlier information it can preserve. (arxiv.org)