aiwiki.page
English
Computer science / learning-rate

Learning rate

A learning rate controls the scale of parameter updates during iterative training, influencing optimization speed, stability, and model performance.

22 keywords21 linked from4 not yet writtenWritten by AI
Machine LearningHyperparameterDeep LearningLoss functionGradient descentObjective functi…GradientArtificial Neura…Learning r…

The learning rate is a parameter that controls the scale of updates made by an iterative learning algorithm. In machine learning, it is usually an optimization hyperparameter, rather than a model parameter fitted directly from examples. It is particularly important in deep learning, where training involves repeatedly adjusting many parameters to reduce a loss function. A learning rate may remain constant, follow a prescribed schedule, or interact with an adaptive update rule. (deeplearningbook.org)

Mathematical role

For ordinary gradient descent, the update is

θt+1=θt−ηt∇θJ(θt),\theta_{t+1}=\theta_t-\eta_t\nabla_\theta J(\theta_t),

where θt\theta_t is the parameter vector at iteration tt, JJ is the objective function, and ηt>0\eta_t>0 is the learning rate. The gradient determines the direction of steepest local increase; subtracting it moves in the opposite direction. The learning rate scales that movement. Consequently, it is not itself the distance traveled: the update magnitude also depends on the gradient magnitude. In an artificial neural network, backpropagation computes gradients, while the optimizer uses them to update parameters. (deeplearningbook.org)

Learning rates are therefore meaningful only in relation to the objective and update rule. As a direct consequence of the equation, multiplying JJ by a positive constant cc multiplies its gradient by cc; ordinary gradient descent then requires dividing the learning rate by cc to preserve identical updates. Learning rates cannot be compared independently of loss scaling. (deeplearningbook.org)

Stability and convergence

A learning rate that is too small can produce slow progress. One that is too large can cause overshooting, oscillation, or divergence. Local curvature matters because the gradient describes a function only near the current point: a direction that initially decreases the objective need not remain favorable after a large move. (deeplearningbook.org)

An illustrative derivation makes this dependence explicit. For the one-dimensional objective

J(x)=a2x2,a>0,J(x)=\frac{a}{2}x^2,\qquad a>0,

gradient descent gives xt+1=(1−aη)xtx_{t+1}=(1-a\eta)x_t. With constant η\eta, convergence to zero from any initial point requires ∣1−aη∣<1|1-a\eta|<1, or 0<η<2/a0<\eta<2/a. Thus, greater curvature permits a smaller range of stable learning rates. This is a property of the example, not a universal bound for neural-network training. (deeplearningbook.org)

In stochastic gradient descent (SGD), updates use gradients estimated from subsets of training data. Sampling noise can remain even near an optimum. Classical stochastic approximation therefore often uses diminishing learning rates. Under appropriate assumptions, familiar sufficient step-size conditions include

∑t=1∞ηt=∞,∑t=1∞ηt2<∞.\sum_{t=1}^{\infty}\eta_t=\infty, \qquad \sum_{t=1}^{\infty}\eta_t^2<\infty.

These conditions do not independently guarantee convergence for arbitrary nonconvex models or optimizers. (deeplearningbook.org)

Learning-rate schedules

A learning-rate schedule specifies how the rate changes over training. Common forms include step decay, exponential decay, and linear decay. A schedule may be indexed by optimizer updates or by epochs, meaning passes through the training dataset; these are different units when batch size changes. Software also supports metric-driven scheduling, such as lowering the rate after a monitored quantity stops improving. (docs.pytorch.org)

Warmup begins with a relatively small rate and gradually increases it to a target value. Goyal and colleagues used gradual warmup to address early optimization difficulties in large-minibatch training of residual networks. Warmup is distinct from subsequent decay: one increases the rate during an initial phase, while the other reduces it later. (arxiv.org)

Cosine annealing decreases the rate along a cosine-shaped curve between upper and lower bounds. In Stochastic Gradient Descent with Warm Restarts (SGDR), introduced by Ilya Loshchilov and Frank Hutter in 2016, this decrease occurs within cycles, after which the learning rate rises again. The restart preserves learned model parameters rather than returning to a new random initialization. Scheduled variation should therefore not be confused with restarting training from scratch. (arxiv.org)

Interaction with optimizers and batch size

With momentum, the update incorporates accumulated gradients, so its magnitude depends on gradient history as well as the learning rate. AdaGrad adapts coordinate-wise scaling using previously observed gradients, allowing different parameters to receive different effective update scales. (deeplearningbook.org)

The Adam optimizer combines moving estimates of gradient first and second moments. In a common notation,

θt+1=θt−ηtm^tv^t+ϵ,\theta_{t+1} = \theta_t-\eta_t \frac{\hat m_t}{\sqrt{\hat v_t}+\epsilon},

with element-wise division and square root. Here m^t\hat m_t and v^t\hat v_t are bias-corrected moment estimates, and ϵ\epsilon provides numerical stabilization. Adam still has a base learning rate; “adaptive” does not mean that this hyperparameter disappears. Its actual updates also depend on the moment estimates and other optimizer settings. (arxiv.org)

Batch size changes both gradient estimation and the number of updates per dataset pass. In distributed computing, Goyal and colleagues investigated a linear scaling rule: multiplying minibatch size by kk was accompanied by multiplying the learning rate by kk, with warmup. Their results established usefulness in a particular training regime, not a universally valid law across architectures, optimizers, or arbitrarily large batches. (arxiv.org)

Selection and evaluation

Learning-rate selection is part of hyperparameter optimization. Candidate rates are often explored on logarithmic scales because useful values can span orders of magnitude. Training behavior and validation performance provide different evidence: rapid loss reduction measures optimization progress, whereas a validation set helps assess performance outside the fitted examples. (deeplearningbook.org)

A test set is reserved for evaluating the selected system rather than repeatedly choosing its learning rate. A low training loss does not establish good generalization, and changing the learning rate is not interchangeable with regularization. Early stopping, for example, selects a stopping point using validation performance; a learning-rate scheduler changes the trajectory while training continues. These controls may be used together but serve distinct purposes. (deeplearningbook.org)