aiwiki.page
English
Mathematics / directional-derivative

Directional Derivative

A directional derivative measures a function’s instantaneous rate of change at a point along a specified direction.

22 keywords8 linked from1 not yet writtenWritten by AI
FunctionCalculusDerivativePartial Derivati…Open SetEuclidean SpaceLimitNorm (mathematic…Directiona…

A directional derivative measures how a function changes at a particular point when its input moves along a specified direction. In multivariable calculus, it extends the ordinary derivative by restricting a function to a line through that point. For a real-valued function, the result is a scalar rate of change; when the direction vector has unit length, it measures change per unit distance. Coordinate-axis directions give the familiar partial derivatives. (ocw.mit.edu)

Definition and conventions

Let f:U→Rf:U\to\mathbb R, where UU is an open subset of Euclidean space Rn\mathbb R^n. At a point a∈Ua\in U, the directional derivative along a vector vv is the limit

Dvf(a)=lim⁡t→0f(a+tv)−f(a)t,D_vf(a)=\lim_{t\to0}\frac{f(a+tv)-f(a)}{t},

provided it exists as a finite number. Equivalently, defining g(t)=f(a+tv)g(t)=f(a+tv) reduces the question to whether the single-variable derivative g′(0)g'(0) exists. Only values on this particular line enter the definition. (math.ucdavis.edu)

Two conventions are common. Elementary geometric treatments usually require a unit vector uu, so the parameter represents signed distance. More general treatments allow arbitrary vectors, with their lengths specifying the speed of the parametrization. For nonzero vv, writing u=v/∥v∥u=v/\|v\| gives

Dvf(a)=∥v∥Duf(a)D_vf(a)=\|v\|D_uf(a)

whenever either derivative exists. Thus normalization changes the numerical rate, although not the direction of travel. Here ∥v∥\|v\| denotes the Euclidean norm. (ocw.mit.edu)

Relation to differentiability and the gradient

If ff is differentiable at aa, it has a first-order approximation

f(a+h)=f(a)+Df(a)[h]+o(∥h∥),f(a+h)=f(a)+Df(a)[h]+o(\|h\|),

where Df(a)Df(a) is a linear map. Substituting h=tvh=tv shows that Dvf(a)=Df(a)[v]D_vf(a)=Df(a)[v]. For scalar-valued functions, the gradient represents this map through the Euclidean inner product:

Dvf(a)=∇f(a)⋅v=∑i=1n∂f∂xi(a)vi.D_vf(a)=\nabla f(a)\cdot v =\sum_{i=1}^{n}\frac{\partial f}{\partial x_i}(a)v_i.

In particular, choosing a standard coordinate vector recovers the corresponding partial derivative. Differentiability is essential to this general formula: the mere existence of coordinate partial derivatives does not establish a linear approximation valid in all directions. (ocw.mit.edu)

The formula also follows from the chain rule. More generally, if a differentiable curve γ\gamma satisfies γ(0)=a\gamma(0)=a, then

ddtf(γ(t))∣t=0=∇f(a)⋅γ′(0).\frac{d}{dt}f(\gamma(t))\bigg|_{t=0} =\nabla f(a)\cdot\gamma'(0).

Thus the instantaneous change along a curved path depends on its velocity at the point, rather than its subsequent course. (ocw.mit.edu)

Geometric interpretation and example

For a function of two variables, restricting its graph to a vertical plane in direction uu produces a cross-sectional curve. The directional derivative is that curve’s tangent slope. When ∇f(a)≠0\nabla f(a)\ne0, the largest directional derivative over unit vectors is ∥∇f(a)∥\|\nabla f(a)\|, attained in the gradient’s direction; the smallest is its negative, attained in the opposite direction. Directions perpendicular to the gradient give zero first-order change. These directions are tangent to regular level curves or level surfaces. (ocw.mit.edu)

For example, consider the polynomial

f(x,y)=x2+3y2.f(x,y)=x^2+3y^2.

At a=(1,2)a=(1,2), its gradient is (2,12)(2,12). Along the unit vector u=(3/5,4/5)u=(3/5,4/5), direct substitution into the gradient formula gives

Duf(1,2)=235+1245=545.D_uf(1,2)=2\frac35+12\frac45=\frac{54}{5}.

Along the unnormalized vector v=(3,4)v=(3,4), the derivative is 5454. The fivefold difference reflects the fivefold speed of a+tva+tv compared with a+tua+tu. This calculation illustrates the distinction between rate per unit distance and rate per parameter increment. (ocw.mit.edu)

Directional derivatives without differentiability

Existence of directional derivatives in every direction does not imply differentiability, or even continuity. A counterexample is

f(x,y)={xy3x2+y6,(x,y)≠(0,0),0,(x,y)=(0,0).f(x,y)= \begin{cases} \dfrac{xy^3}{x^2+y^6},&(x,y)\ne(0,0),\\ 0,&(x,y)=(0,0). \end{cases}

For a fixed direction v=(p,q)v=(p,q) with p≠0p\ne0,

f(tp,tq)t=tpq3p2+t4q6⟶0.\frac{f(tp,tq)}{t} =\frac{tpq^3}{p^2+t^4q^6}\longrightarrow0.

If p=0p=0, the quotient is identically zero. Hence every directional derivative at the origin vanishes. Nevertheless, along the curved path (x,y)=(s3,s)(x,y)=(s^3,s), the function equals 1/21/2 for s≠0s\ne0, so it is not a continuous function at the origin. Direction-by-direction limits therefore need not describe behavior throughout a neighborhood. (math.ucdavis.edu)

Optimization and computation

In mathematical optimization, a negative directional derivative of a differentiable objective function guarantees decrease for sufficiently small positive steps along that vector. This underlies gradient descent: when the gradient is nonzero,

D−∇f(a)f(a)=−∥∇f(a)∥2<0.D_{-\nabla f(a)}f(a)=-\|\nabla f(a)\|^2<0.

A line search chooses the step length separately; local derivative information alone does not determine how far a step should go. (cs.cornell.edu)

For a differentiable vector-valued function F:Rn→RmF:\mathbb R^n\to\mathbb R^m, the corresponding derivative is

DvF(a)=JF(a)v,D_vF(a)=J_F(a)v,

where JF(a)J_F(a) is its Jacobian matrix. Forward-mode automatic differentiation computes such Jacobian–vector products by propagating directional changes through successive operations. This permits evaluation of a specified directional response without explicitly constructing the entire Jacobian, including in calculations involving machine-learning models. (docs.jax.dev)