In calculus, the gradient of a differentiable, scalar-valued function is a vector that describes its first-order change with respect to its input variables. In Euclidean space, its components are the function’s partial derivatives. When nonzero, it points in the direction of greatest local increase, and its magnitude gives the rate of increase in that direction. Written or , the gradient connects differentiation with geometry and optimization. (ocw.mit.edu)
Definition and notation
Let be differentiable on an open set . In Cartesian coordinates with the standard Euclidean inner product, its gradient is
Each partial derivative measures change in one coordinate while holding the others fixed. The symbol , called nabla or del, denotes a differential operator; applying it to a scalar-valued function produces a vector-valued function. Assigning this vector to every point produces a gradient vector field. (ocw.mit.edu)
The gradient represents the function’s derivative through the local expansion
Thus, the dot product predicts the first-order change caused by a small displacement. The remainder becomes negligible relative to the displacement as its length approaches zero. This linear approximation, rather than the mere existence of individual partial derivatives, underlies the gradient’s geometric interpretation. (ocw.mit.edu)
Directional change and level sets
For a unit vector , the directional derivative is
If the gradient is nonzero, this expression is largest when points along the gradient and smallest in the opposite direction. The extreme values are and . Directions perpendicular to the gradient have zero first-order change. “Greatest increase” therefore refers to an instantaneous rate measured per unit Euclidean distance, not necessarily to the best finite displacement. (ocw.mit.edu)
A level set consists of points satisfying . At a regular point, where the gradient is nonzero, it is perpendicular to every tangent direction of that level set. In two dimensions these sets are level curves, represented by contour lines; in three dimensions they are level surfaces. The gradient belongs to the function’s input space, so the gradient of a two-variable function is not itself a three-dimensional normal to its graph. (ocw.mit.edu)
For example, direct differentiation of
gives . At , the gradient is , its magnitude is , and the steepest-ascent unit direction is . It is normal to the ellipse at that point.
Algebraic relationships
Gradients obey linearity and the product rule:
for constants and differentiable scalar functions . For a differentiable map and scalar function , the chain rule becomes
where is the Jacobian matrix and the superscript denotes matrix transpose. This expresses how sensitivities propagate through compositions. For scalar output, the Jacobian is conventionally a row matrix, whereas the gradient is conventionally a column vector. (jmlr.org)
Optimization and computation
In mathematical optimization, gradient descent updates an input or parameter vector by
The positive step size , often called a learning rate, controls the displacement. A line search can select a step that sufficiently decreases the objective. Although the negative gradient supplies a local descent direction when nonzero, convergence requires suitable assumptions and step-size choices. (web.stanford.edu)
At an unconstrained differentiable local minimum in the interior of the domain, the gradient vanishes. The same condition can also occur at a maximum or saddle point, so it does not generally identify a minimum. For a differentiable convex function on an open convex domain, however, a zero gradient identifies a global minimum. (web.stanford.edu)
In machine learning, gradients of a loss function describe sensitivity to model parameters. Automatic differentiation computes derivatives by applying the chain rule to elementary operations. Backpropagation in neural networks is a specialized application of reverse-mode differentiation; it computes gradients rather than selecting the optimization update itself. Finite-difference approximations instead estimate derivatives from nearby function evaluations and involve a trade-off between truncation and rounding errors. (jmlr.org)
Physical and geometric interpretations
In physics, gradients describe spatial variation. A temperature gradient points toward increasing temperature; for an isotropic material obeying Fourier’s law, heat flux is proportional to its negative. In electrostatics, the electric field satisfies , where is electric potential. A spatial gradient has units of the scalar quantity divided by length. (feynmanlectures.caltech.edu)
In differential geometry, the definition depends on a metric. On a Riemannian manifold with metric tensor , the gradient is the unique tangent vector satisfying
for every tangent vector . The differential records directional change independently of a metric, while the metric converts that information into a vector and determines what “steepest” means. Consequently, simply collecting coordinate partial derivatives gives gradient components only in appropriate Euclidean coordinates; general coordinates require metric factors. (math.stanford.edu)