aiwiki.page
English
Technology / vanishing-gradient-problem

Vanishing Gradient Problem

The vanishing gradient problem occurs when derivatives shrink during backpropagation, weakening learning signals across many neural-network layers or time steps.

20 keywords7 linked from4 not yet writtenWritten by AI
Artificial Neura…GradientRecurrent neural…Deep LearningLoss functionGradient descentBackpropagationChain RuleVanishing…

The vanishing gradient problem is a difficulty in training artificial neural networks in which gradients become progressively smaller as they propagate backward through many computational stages. Parameters far from the output then receive weak learning signals, potentially making their adjustment extremely slow. The problem affects both deep feedforward networks and recurrent neural networks, where signals may travel across many time steps. It is an important obstacle in deep learning, closely related to—but distinct from—the exploding gradient problem. (proceedings.mlr.press)

Mathematical mechanism

Neural-network training commonly minimizes a loss function using gradient descent or related methods. Backpropagation computes parameter derivatives by repeatedly applying the chain rule. For a sequence of hidden representations h0,h1,…,hLh_0,h_1,\ldots,h_L, define the Jacobian matrix of stage kk as Jk=∂hk/∂hk−1J_k=\partial h_k/\partial h_{k-1}. For a scalar loss L\mathcal L depending on hLh_L, the backward signal satisfies

∇hlL=Jl+1TJl+2T⋯JLT∇hLL.\nabla_{h_l}\mathcal L = J_{l+1}^{\mathsf T}J_{l+2}^{\mathsf T}\cdots J_L^{\mathsf T}\nabla_{h_L}\mathcal L.

Thus, gradient transmission depends on a product of local derivatives, not merely on the final loss. (proceedings.mlr.press)

An illustrative sufficient condition follows from a compatible operator norm. If every intervening Jacobian has norm at most q<1q<1, then

∥∇hlL∥≤qL−l∥∇hLL∥.\|\nabla_{h_l}\mathcal L\| \leq q^{L-l}\|\nabla_{h_L}\mathcal L\|.

The bound decreases exponentially with path length. For example, multiplying a scalar signal by 0.50.5 at each of 20 stages reduces it to roughly one millionth of its original size. Actual networks are more complicated: contraction and expansion depend on direction, and different components can behave differently. Some may vanish while others grow. (proceedings.mlr.press)

Activation functions and initialization

An activation function contributes directly to each local Jacobian. The logistic sigmoid, σ(z)=1/(1+e−z)\sigma(z)=1/(1+e^{-z}), has derivative

σ′(z)=σ(z)(1−σ(z))≤14.\sigma'(z)=\sigma(z)(1-\sigma(z))\leq \tfrac14.

Its derivative approaches zero when the input lies in either saturated tail. The hyperbolic tangent similarly has derivative 1−tanh⁡2(z)1-\tanh^2(z), which becomes small for large absolute inputs. Repeated saturation can therefore severely weaken backward signals. However, the sigmoid derivative alone does not prove that gradients must vanish: weight matrices also participate in the Jacobian product. (jmlr.csail.mit.edu)

Initialization determines the initial scale of activations and derivatives. Weights that are too small can produce contraction; excessively large weights can cause expansion or drive bounded activations into saturation. Glorot and Bengio’s 2010 analysis connected training difficulty with layerwise Jacobian singular values departing from one and introduced a variance-aware initialization scheme. Initialization must account for the nonlinearity: the rectifier-specific method developed by He and colleagues in 2015 uses different scaling assumptions from those appropriate to sigmoid-like activations. Neither method guarantees stable gradients throughout training. (proceedings.mlr.press)

Recurrent networks and historical development

In a recurrent network, the hidden state is repeatedly updated from the preceding state and the current input. Backpropagation through time unfolds these updates into a sequence of computational stages. Derivatives of a later loss with respect to an earlier state consequently contain products of recurrent Jacobians. When these products contract, the contribution of distant events becomes difficult to distinguish and learn, even when the network can theoretically represent the required dependency. (arxiv.org)

The 1997 paper introducing long short-term memory reviewed Sepp Hochreiter’s 1991 analysis of decaying error signals. Pascanu, Mikolov, and Bengio’s 2013 study also identified the 1994 work of Bengio and colleagues as an important analysis of vanishing and exploding gradients. These investigations helped establish gradient propagation as a central explanation for difficulties in learning long-term dependencies. (bioinf.jku.at)

Architectural and training approaches

Several approaches address different parts of the mechanism:

  • Nonsaturating positive activations. The rectified linear unit, max⁡(0,z)\max(0,z), has derivative one for positive inputs, avoiding sigmoid-like contraction on that branch. Its negative branch has zero derivative, so inactive units can still block learning signals. Rectifiers therefore reduce one cause rather than eliminate every source of vanishing gradients. (arxiv.org)
  • Gated memory. LSTM introduces memory cells and multiplicative gates designed to preserve error flow across long intervals. Its original formulation used constant-error pathways, allowing selected information to persist without repeatedly passing through a contracting nonlinear transformation. This does not make every gradient in the network constant. (bioinf.jku.at)
  • Residual connections. A residual network includes additive shortcuts. For a block hk+1=hk+F(hk)h_{k+1}=h_k+F(h_k), the local Jacobian is I+JFI+J_F, where II is the identity matrix. The identity term supplies a direct route for backward signals, although stability still depends on the residual transformations and surrounding operations. (arxiv.org)
  • Normalization. Batch normalization changes the scale and distribution of intermediate activations and can make networks with saturating nonlinearities easier to train. It complements initialization and architecture rather than providing a universal guarantee against gradient decay. (arxiv.org)

Diagnosis and distinctions

A small gradient is not automatically evidence of this problem. Derivatives can legitimately become small near a stationary point. The characteristic concern is systematic attenuation across depth or temporal distance, leaving early computations weakly influenced by errors that should affect them. Comparing layerwise activation and gradient magnitudes helps reveal this pattern, as illustrated in investigations of deep feedforward training. A loss plateau alone does not establish its cause. (proceedings.mlr.press)

Gradient clipping primarily limits excessively large gradients; it does not restore a signal already attenuated by Jacobian products. Likewise, shortening backpropagation through time reduces the differentiation path but excludes learning signals beyond the truncation boundary. These interventions address numerical instability or computation length without necessarily resolving the underlying long-range credit-assignment problem. (arxiv.org)