DERIVATIVES & SLOPES ~ The Mathematics of Instantaneous Change

Krishnanand
By Krishnanand Chapter 02 · Calculus Series
Hey learner!

In Chapter 01 we saw why humanity needed a new mathematics to handle continuous change. The derivative is the first tool of that new mathematics—a precise instrument for measuring exactly how fast something is changing at a single, frozen instant. Speedometers, stock tickers, neural-network gradients: all of them are derivatives in disguise.

1. The Problem: Average Speed vs. Instantaneous Speed

Suppose you drive 120 kilometers in 2 hours. Your average speed is trivial: \(120/2 = 60\) km/h. But that single number hides everything—the stoplights, the highway accelerations, the moment you slowed for a toll booth. At any given instant you might have been travelling 100 km/h, 20 km/h, or exactly 0 km/h while waiting at a red light.

Average speed is a question about two points in time. Instantaneous speed is a question about a single point—and asking "how far did I travel in zero seconds?" gives the useless answer \(0/0\). Resolving this paradox is what a derivative actually does.

The trick, invented independently by Newton and Leibniz, is to compute the average speed over a small interval and then shrink that interval toward zero. As the interval collapses, the average speed stops wobbling and settles to a single, well-defined value: the derivative.

2. The Formal Limit Definition

Let \(f\) be a function of \(x\). Pick a point \(x\) and a tiny shift \(h\). The average rate of change of \(f\) between \(x\) and \(x+h\) is the slope of the secant line connecting those two points:

$$\text{slope}_{\text{secant}} \;=\; \frac{f(x+h) - f(x)}{h}$$

The derivative is the limit of this expression as the gap \(h\) collapses to zero. If the limit exists, we say \(f\) is differentiable at \(x\), and we write:

$$f'(x) \;=\; \frac{df}{dx} \;=\; \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}$$

This deceptively small formula carries the entire philosophical weight of calculus. The \(h \to 0\) is what allows us to probe a single instant; the fact that we take a limit rather than just plugging in \(h=0\) is what saves us from the \(0/0\) paradox.

\(f'(x)\)

Lagrange Notation

Read "f-prime of x." Compact, used when the independent variable is obvious from context.

\(\dfrac{df}{dx}\)

Leibniz Notation

Read "d-f-d-x." Emphasizes the variable we differentiate against—critical in multivariable calculus and chain rule.

\(\dot{f}\)

Newton's Fluxion

The dot notation, still used in physics for time derivatives: \(\dot{x} = dx/dt\).

3. Geometric Intuition: The Slope of the Tangent Line

Algebraically, the derivative is a limit. Geometrically, it is the slope of the tangent line to the graph of \(f\) at the point \((x, f(x))\). The secant line from the previous section pivots around \((x, f(x))\) as \(h\) shrinks; in the limit, it becomes the tangent—the unique line that just "kisses" the curve at that point without crossing it (assuming the function is smooth enough).

If \(f'(x) > 0\), the function is rising as we move right; if \(f'(x) < 0\), it is falling; if \(f'(x) = 0\), the tangent is perfectly horizontal and we are at a local extremum (a peak or a valley). This is the geometric reason derivatives are the engine of optimization—finding the lowest point of an error curve is equivalent to finding where the derivative is zero.

The equation of the tangent line at a point \(x=a\) follows directly from the point-slope form:

$$y \;=\; f(a) \;+\; f'(a)\,(x - a)$$

This linear approximation, also called the first-order Taylor approximation, is the foundation of Newton's method for root-finding and of every gradient-descent update in machine learning.

4. Differentiation Rules: How to Actually Compute Derivatives

Using the limit definition for every function would be exhausting. Fortunately, calculus builds up a small library of differentiation rules that let us compute derivatives algebraically without re-deriving the limit each time. Each rule below can be proven from the limit definition; together they cover almost every function a working engineer encounters.

The Power Rule

For any real number \(n\) (positive, negative, or fractional):

$$\frac{d}{dx}\, x^n \;=\; n\, x^{n-1}$$

This single rule handles polynomials, square roots, and reciprocals. The derivative of \(x^3\) is \(3x^2\); the derivative of \(\sqrt{x} = x^{1/2}\) is \(\tfrac{1}{2}x^{-1/2}\); the derivative of \(1/x = x^{-1}\) is \(-x^{-2}\).

The Sum Rule

Derivatives distribute over addition:

$$\frac{d}{dx}\,[f(x) + g(x)] \;=\; f'(x) + g'(x)$$

This means each term in a polynomial can be differentiated independently, which is exactly how we attack a long expression: term by term.

The Product Rule

When two functions are multiplied, their derivatives interact in a specific way:

$$\frac{d}{dx}\,[f(x)\,g(x)] \;=\; f'(x)\,g(x) \;+\; f(x)\,g'(x)$$

The mnemonic "first times derivative of second, plus second times derivative of first" is the way most engineers remember it. The product rule is also the gateway to the chain rule covered in Chapter 04.

The Quotient Rule

For a ratio of two functions:

$$\frac{d}{dx}\!\left[\frac{f(x)}{g(x)}\right] \;=\; \frac{f'(x)\,g(x) - f(x)\,g'(x)}{[g(x)]^2}$$

"Low D-high minus high D-low, square the bottom and away we go"—a silly rhyme, but it survives because it works.

Worked Example: Differentiate \(f(x) = 3x^4 + 5x^2 - 7x + 2\)

Apply the sum rule to handle each term independently, then apply the power rule to each term:

\(f'(x) = 4 \cdot 3x^{4-1} + 2 \cdot 5x^{2-1} - 1 \cdot 7x^{1-1} + 0\)

\(f'(x) = 12x^3 + 10x - 7\)

The constant \(2\) differentiates to zero because a flat line has zero slope. The \(-7x\) term differentiates to a constant \(-7\) because \(x^1 \to 1 \cdot x^0 = 1\). This is the rhythm of polynomial differentiation.

Worked Example: Differentiate \(f(x) = x^2 \sin(x)\)

Use the product rule. Let \(u = x^2\) and \(v = \sin(x)\). Then \(u' = 2x\) and \(v' = \cos(x)\).

\(f'(x) = u'v + uv' = 2x\sin(x) + x^2\cos(x)\)

Notice the result is a sum of two terms, not a single product. The product rule always produces two terms; forgetting the second one is the most common mistake students make on exams.

5. The CS / ML Connection: Gradient Descent

This is where derivatives stop being a math-class chore and become the engine of modern artificial intelligence. Every neural network you have ever heard of—GPT, Stable Diffusion, ResNet—is trained using an algorithm called gradient descent, and the "gradient" in that name is literally a multidimensional derivative.

Imagine a hiker trying to walk down a mountain in heavy fog. They cannot see the valley below, but they can feel the slope of the ground beneath their feet. They take a small step in the steepest downward direction, re-evaluate, and step again. After many iterations, they reach the bottom. That is gradient descent: the "mountain" is the loss function, and the "slope" is its derivative.

Mathematically, we have a model with parameters \(\theta\) and a loss function \(L(\theta)\) (often the Mean Squared Error covered in our MSE module). The gradient \(\nabla L\) is the vector of partial derivatives of \(L\) with respect to each parameter. We update the parameters in the negative gradient direction:

$$\theta_{\text{new}} \;=\; \theta_{\text{old}} \;-\; \eta \,\nabla L(\theta_{\text{old}})$$

Here \(\eta\) is the learning rate—a small positive number controlling step size. The derivative tells us which way is down; the learning rate tells us how big a step to take. Set \(\eta\) too large and you overshoot the valley; set it too small and training takes forever. Derivatives are not optional knowledge for a machine-learning engineer—they are the literal math that powers the field.

Visualizing the Derivative

Grant Sanderson's Essence of Calculus series is the gold standard for building intuition about derivatives. Watch him derive the limit definition from first principles—it is the cleanest explanation of why the derivative "works" that we know of.

Key Takeaway: The derivative is not a magic formula. It is the answer to a precise question: "if I shrink the gap between two points all the way to zero, what value does the average rate of change settle on?"

← Previous: History of Calculus Next: Integrals & Area Under Curves →