THE CHAIN RULE ~ The Mathematics of Composition

Krishnanand
By Krishnanand Chapter 04 · Calculus Series · Final Chapter
Hey learner!

In Chapter 02, you learned to differentiate simple functions like \(x^2\), \(\sin(x)\), and \(e^x\). But the real world rarely hands you simple functions. It hands you compositions: \(\sin(x^2)\), \(e^{-x^2/2}\), \(\log(1 + e^{w \cdot x})\). These are functions nested inside other functions. The chain rule is the universal tool for differentiating them—and, not coincidentally, the mathematical engine behind every backpropagation step in every neural network on Earth.

1. Function Composition: Functions Inside Functions

When we write \(f(g(x))\), we are composing two functions: the outer function \(f\) takes the output of the inner function \(g\) as its input. The notation \((f \circ g)(x)\) means the same thing. Composition is the natural way to build complex systems out of simple parts—each layer does one job, and the next layer takes the previous output and refines it further.

For example, \(\sin(x^2)\) is the composition of "square" inside "sine." To evaluate it for \(x = 3\), we first square 3 (giving 9), then take the sine of 9. The order matters: \(f(g(x))\) is generally not equal to \(g(f(x))\). This is why a deep neural network with the layers in a different order produces a completely different function.

Every deep neural network is, mathematically, a single giant composition: \(\text{output} = f_L(f_{L-1}(\dots f_2(f_1(\text{input})) \dots))\). The depth of the network is the depth of the composition. To train this network—to figure out how each weight should change to reduce the loss—we need to differentiate the entire composition with respect to every weight. The chain rule is what makes that mathematically possible.

2. The Chain Rule Formula

Let \(y = f(u)\) and \(u = g(x)\), so \(y = f(g(x))\). The chain rule says:

$$\frac{dy}{dx} \;=\; \frac{dy}{du} \;\cdot\; \frac{du}{dx}$$

In words: the rate of change of \(y\) with respect to \(x\) equals the rate of change of \(y\) with respect to the intermediate variable \(u\), multiplied by the rate of change of \(u\) with respect to \(x\). The Leibniz notation makes this look almost like fraction cancellation—and that intuition (treating \(dy/du\) and \(du/dx\) as if they were fractions whose \(du\)'s cancel) is a useful mnemonic, though the actual proof rests on the limit definition.

In Lagrange's prime notation, the same rule reads:

$$(f \circ g)'(x) \;=\; f'(g(x)) \;\cdot\; g'(x)$$

The crucial subtlety is that \(f'\) is evaluated at \(g(x)\), not at \(x\). The outer function's derivative is computed at the inner function's value. Forgetting this is the second most common student mistake (right after forgetting the product rule's second term).

Worked Example: Differentiate \(y = \sin(x^2)\)

Identify the composition: outer \(f(u) = \sin(u)\), inner \(u = g(x) = x^2\).

Compute each derivative separately: \(f'(u) = \cos(u)\) and \(g'(x) = 2x\).

Apply the chain rule: \(\frac{dy}{dx} = \cos(u) \cdot 2x = 2x \cos(x^2)\).

Substitute back \(u = x^2\) to get the final answer: \(\boxed{2x \cos(x^2)}\). Notice the inner function \(x^2\) is preserved inside the cosine—this is the hallmark of a chain-rule derivative.

Worked Example: Differentiate \(y = e^{-x^2/2}\)

This is the Gaussian function—the bell curve that defines probability theory. It is a triple composition: \(e^{(\cdot)}\) wrapping \(-(\cdot)/2\) wrapping \(x^2\).

Outer: \(f(u) = e^u\), so \(f'(u) = e^u\). Middle: \(u = g(v) = -v/2\), so \(g'(v) = -1/2\). Inner: \(v = h(x) = x^2\), so \(h'(x) = 2x\).

Chain rule (three layers): \(\frac{dy}{dx} = f'(u) \cdot g'(v) \cdot h'(x) = e^u \cdot \left(-\tfrac{1}{2}\right) \cdot 2x\).

Simplify: \(-x \, e^{-x^2/2}\). This is the derivative of the Gaussian. Notice how the chain rule, applied layer by layer, naturally produces the polynomial factor in front of the exponential—a feature that recurs throughout statistics and physics.

3. The Multi-Variable Chain Rule

In machine learning, we rarely deal with single-variable functions. A neural network's loss depends on hundreds of millions of parameters simultaneously. The chain rule generalizes naturally to multiple variables via the total derivative:

If \(y = f(u_1, u_2, \dots, u_n)\) and each \(u_i = g_i(x)\), then:

$$\frac{dy}{dx} \;=\; \sum_{i=1}^{n} \frac{\partial f}{\partial u_i} \cdot \frac{du_i}{dx}$$

The symbol \(\partial\) denotes a partial derivative—the derivative of \(f\) with respect to \(u_i\) holding all other \(u\)'s constant. The total derivative is the sum of contributions flowing through each "wire" in the computational graph. This sum-over-paths formulation is exactly what automatic differentiation software (PyTorch's autograd, TensorFlow's GradientTape) implements under the hood.

\(\partial\)

Partial Derivative

Rate of change of a multivariable function with respect to one variable, holding the others fixed.

\(\nabla\)

Gradient

The vector of all partial derivatives of a scalar-valued function. Points in the direction of steepest ascent.

\(J\)

Jacobian

The matrix of all first-order partial derivatives of a vector-valued function. The chain rule for vector functions is matrix multiplication.

4. Backpropagation: The Chain Rule in Action

Backpropagation is not a separate algorithm from calculus—it is the multi-variable chain rule, organized cleverly. The trick that makes backprop efficient is computing each partial derivative once and reusing it for all downstream computations.

Consider a 3-layer neural network: input \(x\), hidden \(h = \sigma(W_1 x + b_1)\), output \(\hat{y} = W_2 h + b_2\), and loss \(L = (y - \hat{y})^2\). To update \(W_1\), we need \(\partial L / \partial W_1\). Unrolling the chain rule by hand:

$$\frac{\partial L}{\partial W_1} \;=\; \frac{\partial L}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial h} \cdot \frac{\partial h}{\partial W_1}$$

Each factor is a small, locally-computable derivative. We multiply them in sequence from the output back to the input. Critically, the same intermediate factors—like \(\partial \hat{y} / \partial h\)—are reused for computing \(\partial L / \partial b_1\) as well. This sharing of intermediate results is what gives backprop its name: the gradient propagates backward through the network, accumulating at each weight.

Without the chain rule, training a deep network would require recomputing every partial from scratch for every weight—an \(O(n^2)\) cost in the number of parameters. With backprop, the entire gradient can be computed in a single forward pass and a single backward pass, costing \(O(n)\). This algorithmic speedup is the difference between "neural networks train in hours" and "neural networks train in years."

Worked Example: Backprop on a 1-Neuron Network

Suppose \(\hat{y} = \sigma(w \cdot x + b)\) with \(x = 2\), \(w = 3\), \(b = 1\), true label \(y = 1\), and \(\sigma\) is the sigmoid. Forward pass: \(z = 3 \cdot 2 + 1 = 7\), \(\hat{y} = \sigma(7) \approx 0.9991\). Loss: \(L = (1 - 0.9991)^2 \approx 8 \times 10^{-7}\).

Backward pass (chain rule): \(\partial L / \partial \hat{y} = -2(1 - 0.9991) \approx -0.0018\). \(\partial \hat{y} / \partial z = \hat{y}(1 - \hat{y}) \approx 0.0009\). \(\partial z / \partial w = x = 2\). Chaining: \(\partial L / \partial w = -0.0018 \times 0.0009 \times 2 \approx -3.2 \times 10^{-6}\).

A single gradient descent step with learning rate \(\eta = 1.0\) would update \(w \leftarrow w - \eta \cdot \partial L / \partial w = 3.0000032\). The weight nudges ever so slightly toward reducing the loss. Multiply this tiny step by billions of iterations across millions of parameters, and you have a trained neural network.

Visualizing the Chain Rule

3Blue1Brown's Essence of Calculus chapter on the chain rule is the cleanest geometric explanation we know of—especially the part about how the "stretch factors" of nested functions multiply.

Key Takeaway: When functions nest, their derivatives multiply. This deceptively simple fact is the mathematical engine of backpropagation—and therefore of every deep learning model ever trained.

🏆 You have completed the Calculus Series.

Four chapters. From the historical paradox of "instantaneous speed" to the chain rule that powers backpropagation in modern neural networks. You now have the mathematical foundation needed to study gradient descent, optimization, and deep learning at the level of derivations rather than recipes.

Next stops: our modules on MSE & Loss Functions, SVD, and Laplacian Smoothing apply these concepts directly.

← Previous: Integrals & Area Under Curves Back to Course Catalogue