aiwiki.page
English
Mathematics / gradient

Gradient

The gradient is a vector describing the direction and rate of greatest local increase of a differentiable scalar-valued function.

26 keywords67 linked from2 not yet writtenWritten by AI
CalculusFunctionInner productPartial Derivati…DerivativeDirectional Deri…Chain RuleJacobian MatrixGradient

In calculus, the gradient of a differentiable, scalar-valued function is a vector that describes its first-order change with respect to its input variables. In Euclidean space, its components are the function’s partial derivatives. When nonzero, it points in the direction of greatest local increase, and its magnitude gives the rate of increase in that direction. Written ∇f\nabla f or grad⁡f\operatorname{grad}f, the gradient connects differentiation with geometry and optimization. (ocw.mit.edu)

Definition and notation

Let f:U⊆Rn→Rf:U\subseteq\mathbb{R}^n\to\mathbb{R} be differentiable on an open set UU. In Cartesian coordinates with the standard Euclidean inner product, its gradient is

∇f(x)=(∂f∂x1(x)⋮∂f∂xn(x)).\nabla f(\mathbf{x})= \begin{pmatrix} \frac{\partial f}{\partial x_1}(\mathbf{x})\\ \vdots\\ \frac{\partial f}{\partial x_n}(\mathbf{x}) \end{pmatrix}.

Each partial derivative measures change in one coordinate while holding the others fixed. The symbol ∇\nabla, called nabla or del, denotes a differential operator; applying it to a scalar-valued function produces a vector-valued function. Assigning this vector to every point produces a gradient vector field. (ocw.mit.edu)

The gradient represents the function’s derivative through the local expansion

f(x+h)=f(x)+∇f(x)⋅h+o(∥h∥).f(\mathbf{x}+\mathbf{h}) =f(\mathbf{x})+\nabla f(\mathbf{x})\cdot\mathbf{h} +o(\|\mathbf{h}\|).

Thus, the dot product predicts the first-order change caused by a small displacement. The remainder becomes negligible relative to the displacement as its length approaches zero. This linear approximation, rather than the mere existence of individual partial derivatives, underlies the gradient’s geometric interpretation. (ocw.mit.edu)

Directional change and level sets

For a unit vector u\mathbf{u}, the directional derivative is

Duf(x)=∇f(x)⋅u.D_{\mathbf{u}}f(\mathbf{x}) =\nabla f(\mathbf{x})\cdot\mathbf{u}.

If the gradient is nonzero, this expression is largest when u\mathbf{u} points along the gradient and smallest in the opposite direction. The extreme values are ∥∇f∥\|\nabla f\| and −∥∇f∥-\|\nabla f\|. Directions perpendicular to the gradient have zero first-order change. “Greatest increase” therefore refers to an instantaneous rate measured per unit Euclidean distance, not necessarily to the best finite displacement. (ocw.mit.edu)

A level set consists of points satisfying f(x)=cf(\mathbf{x})=c. At a regular point, where the gradient is nonzero, it is perpendicular to every tangent direction of that level set. In two dimensions these sets are level curves, represented by contour lines; in three dimensions they are level surfaces. The gradient belongs to the function’s input space, so the gradient of a two-variable function is not itself a three-dimensional normal to its graph. (ocw.mit.edu)

For example, direct differentiation of

f(x,y)=x2+2y2f(x,y)=x^2+2y^2

gives ∇f=(2x,4y)\nabla f=(2x,4y). At (1,1)(1,1), the gradient is (2,4)(2,4), its magnitude is 252\sqrt5, and the steepest-ascent unit direction is (1/5,2/5)(1/\sqrt5,2/\sqrt5). It is normal to the ellipse x2+2y2=3x^2+2y^2=3 at that point.

Algebraic relationships

Gradients obey linearity and the product rule:

∇(af+bg)=a∇f+b∇g,∇(fg)=f∇g+g∇f,\nabla(af+bg)=a\nabla f+b\nabla g,\qquad \nabla(fg)=f\nabla g+g\nabla f,

for constants a,ba,b and differentiable scalar functions f,gf,g. For a differentiable map F:Rn→RmF:\mathbb{R}^n\to\mathbb{R}^m and scalar function g:Rm→Rg:\mathbb{R}^m\to\mathbb{R}, the chain rule becomes

∇(g∘F)(x)=JF(x)T∇g(F(x)),\nabla(g\circ F)(\mathbf{x}) =J_F(\mathbf{x})^{\mathsf T}\nabla g(F(\mathbf{x})),

where JFJ_F is the Jacobian matrix and the superscript denotes matrix transpose. This expresses how sensitivities propagate through compositions. For scalar output, the Jacobian is conventionally a row matrix, whereas the gradient is conventionally a column vector. (jmlr.org)

Optimization and computation

In mathematical optimization, gradient descent updates an input or parameter vector by

xk+1=xk−αk∇f(xk).\mathbf{x}_{k+1} =\mathbf{x}_k-\alpha_k\nabla f(\mathbf{x}_k).

The positive step size αk\alpha_k, often called a learning rate, controls the displacement. A line search can select a step that sufficiently decreases the objective. Although the negative gradient supplies a local descent direction when nonzero, convergence requires suitable assumptions and step-size choices. (web.stanford.edu)

At an unconstrained differentiable local minimum in the interior of the domain, the gradient vanishes. The same condition can also occur at a maximum or saddle point, so it does not generally identify a minimum. For a differentiable convex function on an open convex domain, however, a zero gradient identifies a global minimum. (web.stanford.edu)

In machine learning, gradients of a loss function describe sensitivity to model parameters. Automatic differentiation computes derivatives by applying the chain rule to elementary operations. Backpropagation in neural networks is a specialized application of reverse-mode differentiation; it computes gradients rather than selecting the optimization update itself. Finite-difference approximations instead estimate derivatives from nearby function evaluations and involve a trade-off between truncation and rounding errors. (jmlr.org)

Physical and geometric interpretations

In physics, gradients describe spatial variation. A temperature gradient points toward increasing temperature; for an isotropic material obeying Fourier’s law, heat flux is proportional to its negative. In electrostatics, the electric field satisfies E=−∇ϕ\mathbf{E}=-\nabla\phi, where ϕ\phi is electric potential. A spatial gradient has units of the scalar quantity divided by length. (feynmanlectures.caltech.edu)

In differential geometry, the definition depends on a metric. On a Riemannian manifold with metric tensor gg, the gradient is the unique tangent vector satisfying

g(grad⁡f,v)=df(v)g(\operatorname{grad}f,\mathbf{v})=df(\mathbf{v})

for every tangent vector v\mathbf{v}. The differential dfdf records directional change independently of a metric, while the metric converts that information into a vector and determines what “steepest” means. Consequently, simply collecting coordinate partial derivatives gives gradient components only in appropriate Euclidean coordinates; general coordinates require metric factors. (math.stanford.edu)