In Chapter 02, we used derivatives to answer the question "how fast is something changing at a single instant?" The integral is the inverse operation. It answers a different question: "if I know the rate of change at every instant, how much has accumulated in total?" Distance is the integral of velocity. Energy is the integral of power. Probability is the integral of a density. Every "running total" in physics, finance, and machine learning is an integral in disguise.
Suppose you have a velocity function \(v(t)\) describing how fast a car is moving at every instant \(t\). The derivative question is: "given a position function, find velocity." The integral question runs the other way: "given velocity, find position." Geometrically, the total distance travelled between time \(a\) and time \(b\) is exactly the area trapped between the velocity curve and the horizontal axis over that interval.
This geometric fact is what makes integrals so universally useful. Wherever you see accumulation—water filling a tank at a variable rate, charge building up on a capacitor, money accruing in a variable-interest bank account—the total accumulated amount is the area under the rate curve. The integral is the formal name for that area, computed even when the curve is too complex for simple geometry.
To compute the area under a curve, we apply the same trick derivatives used: take a discrete approximation, then shrink the gaps. Divide the interval \([a, b]\) into \(n\) sub-intervals of equal width \(\Delta x = (b-a)/n\). On each sub-interval, pick a sample point \(x_i^*\) and approximate the area under the curve on that sub-interval with the rectangle of width \(\Delta x\) and height \(f(x_i^*)\).
Summing all these rectangles gives a Riemann sum:
$$S_n \;=\; \sum_{i=1}^{n} f(x_i^*)\, \Delta x$$As \(n \to \infty\), the rectangles become thinner and thinner, and their total area converges to the true area under the curve. The definite integral is defined as the limit of this Riemann sum:
$$\int_{a}^{b} f(x)\, dx \;=\; \lim_{n \to \infty} \sum_{i=1}^{n} f(x_i^*)\, \Delta x$$The elongated \(\int\) symbol is Leibniz's stylized "S" for Summa—a reminder that integration is, at heart, the limit of an infinite sum of infinitely thin slices.
The interval endpoints. We are integrating "from a to b."
An infinitesimally thin slice of width along the x-axis. Also signals which variable we integrate against.
The function we are integrating—the curve whose area we are computing.
The definite integral, written \(\int_a^b f(x)\, dx\), returns a number: the net area between the curve and the x-axis from \(a\) to \(b\). Areas above the axis are positive; areas below are negative. This signed area is what makes the integral a true algebraic object rather than just a geometric one.
The indefinite integral, written without bounds as \(\int f(x)\, dx\), returns a function: the family of all antiderivatives of \(f\). If \(F'(x) = f(x)\), then:
$$\int f(x)\, dx \;=\; F(x) + C$$The constant \(C\) reflects the fact that differentiation kills constants—the derivative of \(F(x) + 5\) is the same as the derivative of \(F(x) + 7\). Therefore many different functions share the same derivative, and the indefinite integral must capture all of them with the \(+ C\) placeholder.
Step 1: Find the antiderivative. The power rule for integration inverts the differentiation power rule: \(\int x^n\, dx = \frac{x^{n+1}}{n+1} + C\) for \(n \neq -1\).
\(\int 3x^2\, dx = 3 \cdot \frac{x^3}{3} + C = x^3 + C\)
Step 2: Apply the limits using the Fundamental Theorem (next section): \(\int_0^{2} 3x^2\, dx = \big[x^3\big]_0^2 = (2)^3 - (0)^3 = 8\).
The area under the curve \(y = 3x^2\) from 0 to 2 is exactly 8 square units.
The single deepest theorem in calculus—and arguably in all of mathematics—is the statement that differentiation and integration are inverse operations. The theorem has two halves.
If we define a new function \(F(x) = \int_a^x f(t)\, dt\), then:
$$F'(x) \;=\; f(x)$$Integrating a function and then differentiating the result recovers the original function. This is the formal sense in which integration is the inverse of differentiation.
If \(F\) is any antiderivative of \(f\) (so \(F' = f\)), then:
$$\int_a^b f(x)\, dx \;=\; F(b) - F(a)$$This is the practical engine of integral calculus. Instead of computing Riemann sums for every integral—an exhausting process—we just find an antiderivative and plug in the limits. Every homework problem, every physics derivation, every ML probability computation leans on this exact shortcut.
Where do integrals hide inside machine learning? Three places, each important.
A continuous probability density function \(p(x)\) does not give probabilities directly. The probability of a single point is zero (a point has no width). To get the probability that \(X\) falls in an interval \([a, b]\), we integrate:
$$P(a \le X \le b) \;=\; \int_a^b p(x)\, dx$$The total area under a probability density must equal 1: \(\int_{-\infty}^{\infty} p(x)\, dx = 1\). This is why the Gaussian, the uniform, the exponential—every probability distribution you have ever met—is normalized by an integral.
The expected value of a function \(g(X)\) under a distribution \(p(x)\) is defined as an integral:
$$\mathbb{E}[g(X)] \;=\; \int_{-\infty}^{\infty} g(x)\, p(x)\, dx$$When you minimize Mean Squared Error (covered in our MSE module) over a continuous distribution, you are literally minimizing an integral—specifically, \(\mathbb{E}[(Y - \hat{Y})^2]\). The MSE formula you know, \(\frac{1}{n}\sum_i (y_i - \hat{y}_i)^2\), is just the discrete version of this integral.
Modern architectures like Neural ODEs treat the depth of a neural network as a continuous variable. Instead of a discrete stack of layers \(h_{t+1} = f(h_t, \theta_t)\), they model the hidden state as a continuous flow: \(\frac{dh}{dt} = f(h(t), t, \theta)\). The forward pass is the solution of an ordinary differential equation, computed by integrating \(f\) from \(t=0\) to \(t=T\). The backward pass uses the adjoint method—a sensitivity equation derived from the FTC.
So when you read that Neural ODEs achieved state-of-the-art results on irregular time-series data, you are watching a 300-year-old theorem about accumulation power a cutting-edge deep learning architecture.
Grant Sanderson returns with the cleanest geometric explanation of the Fundamental Theorem of Calculus you will find anywhere. Watch how he makes the connection between "area under a curve" and "antiderivative" feel obvious.
Key Takeaway: The integral is the inverse of the derivative—not by accident, but because accumulation and rate-of-change are two sides of the same mathematical coin.