The vanishing gradient problem is a difficulty in training artificial neural networks in which gradients become progressively smaller as they propagate backward through many computational stages. Parameters far from the output then receive weak learning signals, potentially making their adjustment extremely slow. The problem affects both deep feedforward networks and recurrent neural networks, where signals may travel across many time steps. It is an important obstacle in deep learning, closely related to—but distinct from—the exploding gradient problem. (proceedings.mlr.press)
Mathematical mechanism
Neural-network training commonly minimizes a loss function using gradient descent or related methods. Backpropagation computes parameter derivatives by repeatedly applying the chain rule. For a sequence of hidden representations , define the Jacobian matrix of stage as . For a scalar loss depending on , the backward signal satisfies
Thus, gradient transmission depends on a product of local derivatives, not merely on the final loss. (proceedings.mlr.press)
An illustrative sufficient condition follows from a compatible operator norm. If every intervening Jacobian has norm at most , then
The bound decreases exponentially with path length. For example, multiplying a scalar signal by at each of 20 stages reduces it to roughly one millionth of its original size. Actual networks are more complicated: contraction and expansion depend on direction, and different components can behave differently. Some may vanish while others grow. (proceedings.mlr.press)
Activation functions and initialization
An activation function contributes directly to each local Jacobian. The logistic sigmoid, , has derivative
Its derivative approaches zero when the input lies in either saturated tail. The hyperbolic tangent similarly has derivative , which becomes small for large absolute inputs. Repeated saturation can therefore severely weaken backward signals. However, the sigmoid derivative alone does not prove that gradients must vanish: weight matrices also participate in the Jacobian product. (jmlr.csail.mit.edu)
Initialization determines the initial scale of activations and derivatives. Weights that are too small can produce contraction; excessively large weights can cause expansion or drive bounded activations into saturation. Glorot and Bengio’s 2010 analysis connected training difficulty with layerwise Jacobian singular values departing from one and introduced a variance-aware initialization scheme. Initialization must account for the nonlinearity: the rectifier-specific method developed by He and colleagues in 2015 uses different scaling assumptions from those appropriate to sigmoid-like activations. Neither method guarantees stable gradients throughout training. (proceedings.mlr.press)
Recurrent networks and historical development
In a recurrent network, the hidden state is repeatedly updated from the preceding state and the current input. Backpropagation through time unfolds these updates into a sequence of computational stages. Derivatives of a later loss with respect to an earlier state consequently contain products of recurrent Jacobians. When these products contract, the contribution of distant events becomes difficult to distinguish and learn, even when the network can theoretically represent the required dependency. (arxiv.org)
The 1997 paper introducing long short-term memory reviewed Sepp Hochreiter’s 1991 analysis of decaying error signals. Pascanu, Mikolov, and Bengio’s 2013 study also identified the 1994 work of Bengio and colleagues as an important analysis of vanishing and exploding gradients. These investigations helped establish gradient propagation as a central explanation for difficulties in learning long-term dependencies. (bioinf.jku.at)
Architectural and training approaches
Several approaches address different parts of the mechanism:
- Nonsaturating positive activations. The rectified linear unit, , has derivative one for positive inputs, avoiding sigmoid-like contraction on that branch. Its negative branch has zero derivative, so inactive units can still block learning signals. Rectifiers therefore reduce one cause rather than eliminate every source of vanishing gradients. (arxiv.org)
- Gated memory. LSTM introduces memory cells and multiplicative gates designed to preserve error flow across long intervals. Its original formulation used constant-error pathways, allowing selected information to persist without repeatedly passing through a contracting nonlinear transformation. This does not make every gradient in the network constant. (bioinf.jku.at)
- Residual connections. A residual network includes additive shortcuts. For a block , the local Jacobian is , where is the identity matrix. The identity term supplies a direct route for backward signals, although stability still depends on the residual transformations and surrounding operations. (arxiv.org)
- Normalization. Batch normalization changes the scale and distribution of intermediate activations and can make networks with saturating nonlinearities easier to train. It complements initialization and architecture rather than providing a universal guarantee against gradient decay. (arxiv.org)
Diagnosis and distinctions
A small gradient is not automatically evidence of this problem. Derivatives can legitimately become small near a stationary point. The characteristic concern is systematic attenuation across depth or temporal distance, leaving early computations weakly influenced by errors that should affect them. Comparing layerwise activation and gradient magnitudes helps reveal this pattern, as illustrated in investigations of deep feedforward training. A loss plateau alone does not establish its cause. (proceedings.mlr.press)
Gradient clipping primarily limits excessively large gradients; it does not restore a signal already attenuated by Jacobian products. Likewise, shortening backpropagation through time reduces the differentiation path but excludes learning signals beyond the truncation boundary. These interventions address numerical instability or computation length without necessarily resolving the underlying long-range credit-assignment problem. (arxiv.org)