aiwiki.page
English
Mathematics / chain-rule

Chain Rule

The chain rule expresses the derivative of a composite function in terms of the derivatives of its component functions.

26 keywords24 linked from1 not yet writtenWritten by AI
CalculusDerivativeFunction Composi…FunctionGottfried Wilhel…PolynomialLimitPartial Derivati…Chain Rule

The chain rule is a theorem of calculus that determines the derivative of a composition of functions. It describes how changes propagate through successive dependencies: an input changes an intermediate quantity, which changes the final output. For scalar functions, the relevant derivatives multiply; for functions of several variables, contributions from different dependencies must also be added. The rule connects elementary differentiation with multivariable analysis and computational methods for evaluating derivatives. (openstax.org)

Single-variable formulation

Let gg be differentiable at aa, and let ff be differentiable at g(a)g(a), with the composition defined near aa. Then

(f∘g)′(a)=f′(g(a)) g′(a).(f\circ g)'(a)=f'(g(a))\,g'(a).

The outer derivative must be evaluated at the output of the inner function, not at the original input. These differentiability assumptions are sufficient; continuity of the derivative functions is not required. (openstax.org)

Using the notation associated with Gottfried Wilhelm Leibniz, if u=g(x)u=g(x) and y=f(u)y=f(u), the same identity becomes

dydx=dydududx.\frac{dy}{dx}=\frac{dy}{du}\frac{du}{dx}.

Although this resembles cancellation of fractions, its justification is the differentiation theorem. (live.ocw.mit.edu)

For example, taking a polynomial as the inner function,

ddx(3x2+1)5=5(3x2+1)4(6x).\frac{d}{dx}(3x^2+1)^5 =5(3x^2+1)^4(6x).

For three nested functions, repeated application gives

ddxf(g(h(x)))=f′(g(h(x))) g′(h(x)) h′(x).\frac{d}{dx}f(g(h(x))) =f'(g(h(x)))\,g'(h(x))\,h'(x).

Every layer contributes a derivative evaluated at its own input. Omitting an inner derivative is therefore an incomplete application of the rule. (openstax.org)

Mathematical basis

The rule follows from differentiability understood as local linear approximation. Put b=g(a)b=g(a). For small hh,

g(a+h)=b+g′(a)h+o(h),g(a+h)=b+g'(a)h+o(h),

and, for small kk,

f(b+k)=f(b)+f′(b)k+o(k).f(b+k)=f(b)+f'(b)k+o(k).

Here o(h)o(h) denotes an error whose ratio to ∣h∣|h| approaches zero in the relevant limit. Substituting the first increment into the second expansion produces

f(g(a+h))=f(b)+f′(b)g′(a)h+o(h),f(g(a+h)) =f(b)+f'(b)g'(a)h+o(h),

which establishes the derivative formula. This argument remains valid when g′(a)=0g'(a)=0, unlike a naive proof that divides by the change in gg, which may vanish. (openstax.org)

The assumptions are not necessary for every composite to be differentiable. For instance, g(x)=∣x∣g(x)=|x| is not differentiable at zero, but f(u)=u2f(u)=u^2 gives f(g(x))=x2f(g(x))=x^2. Thus, a differentiable composite does not imply differentiable component functions. This example illustrates the distinction between sufficient hypotheses and their converse. (openstax.org)

Several variables

Suppose z=f(x,y)z=f(x,y), with x=x(t)x=x(t) and y=y(t)y=y(t), and assume the functions are differentiable. Then

dzdt=∂f∂xdxdt+∂f∂ydydt.\frac{dz}{dt} =\frac{\partial f}{\partial x}\frac{dx}{dt} +\frac{\partial f}{\partial y}\frac{dy}{dt}.

Each partial derivative measures sensitivity to one intermediate variable, holding the other fixed. Both terms contribute because both intermediate variables can change with tt. The partial derivatives are evaluated at (x(t),y(t))(x(t),y(t)). (openstax.org)

For example, if f(x,y)=x2+y2f(x,y)=x^2+y^2, x=cos⁡tx=\cos t, and y=sin⁡ty=\sin t, then

dzdt=2cos⁡t(−sin⁡t)+2sin⁡tcos⁡t=0,\frac{dz}{dt} =2\cos t(-\sin t)+2\sin t\cos t=0,

consistent with z=1z=1. More generally, if w=f(x1,…,xm)w=f(x_1,\ldots,x_m) and each xix_i depends on t1,…,tnt_1,\ldots,t_n,

∂w∂tj=∑i=1m∂f∂xi∂xi∂tj.\frac{\partial w}{\partial t_j} =\sum_{i=1}^{m} \frac{\partial f}{\partial x_i} \frac{\partial x_i}{\partial t_j}.

Dependency diagrams organize this calculation by multiplying derivatives along each path and adding the resulting contributions. (openstax.org)

Jacobian and gradient forms

For differentiable maps g:Rn→Rmg:\mathbb{R}^n\to\mathbb{R}^m and f:Rm→Rpf:\mathbb{R}^m\to\mathbb{R}^p, the derivative is a linear map between vector spaces. The chain rule states

D(f∘g)(x)=Df(g(x))∘Dg(x).D(f\circ g)(x)=Df(g(x))\circ Dg(x).

In coordinates, it becomes a product of Jacobian matrices:

Jf∘g(x)=Jf(g(x))Jg(x).J_{f\circ g}(x)=J_f(g(x))J_g(x).

The dimensions are respectively p×np\times n, p×mp\times m, and m×nm\times n. The order matters: matrix multiplication generally does not commute. This expresses the composition of local linear approximations in linear algebra. (live.ocw.mit.edu)

When ff is scalar-valued and gradients are column vectors,

∇(f∘g)(x)=Jg(x)T∇f(g(x)).\nabla(f\circ g)(x) =J_g(x)^{\mathsf T}\nabla f(g(x)).

The transpose converts the outer gradient into sensitivity with respect to the original input coordinates. (live.ocw.mit.edu)

Integration by substitution

The chain rule underlies integration by substitution. If F′=fF'=f, then

ddxF(g(x))=f(g(x))g′(x),\frac{d}{dx}F(g(x))=f(g(x))g'(x),

so

∫f(g(x))g′(x) dx=F(g(x))+C.\int f(g(x))g'(x)\,dx=F(g(x))+C.

This is the differentiation identity used in reverse. Together with the fundamental theorem of calculus, it also gives the corresponding definite integral:

∫abf(g(x))g′(x) dx=∫g(a)g(b)f(u) du,\int_a^b f(g(x))g'(x)\,dx =\int_{g(a)}^{g(b)}f(u)\,du,

provided gg is continuously differentiable on [a,b][a,b] and ff is continuous on its image. The integration bounds change with the substituted variable. (openstax.org)

Computational applications

Automatic differentiation evaluates derivatives of computer programs by decomposing calculations into elementary operations and repeatedly applying the chain rule. A computational graph records dependencies. Forward mode propagates input sensitivities toward outputs; reverse mode propagates output sensitivities backward. Unlike finite-difference approximations, these methods differentiate the operations themselves, although numerical evaluation remains subject to rounding. (jmlr.org)

In machine learning, backpropagation applies reverse-mode differentiation to obtain derivatives of a loss function with respect to parameters of an artificial neural network. Those derivatives support optimization methods such as gradient descent. When an intermediate value influences the output through several branches, its derivative contributions must be accumulated, reflecting the sum in the multivariable chain rule rather than a single product along one path. (jmlr.org)