AI and machine learning

The calculus behind modern AI: loss functions, gradients, backpropagation, optimizers and probabilistic models.

21 topics

The calculus behind modern AI

Strip a neural network of its jargon and what remains is calculus: a big differentiable function, a scalar loss that measures its error, and an algorithm that follows the derivative of the loss downhill.

Calculus
   ├── Derivatives ──► Gradients ──► Optimization ──► Gradient descent
   ├── Partial derivatives + chain rule ──► Backpropagation
   ├── Integrals ──► Probability ──► Probabilistic ML (likelihoods, Bayes, diffusion)
   └── Linear algebra + calculus ──► Neural networks

One training step, end to end:

input x ─► layer z = Wx + b ─► activation a = σ(z) ─► prediction ŷ ─► loss L(ŷ, y)
                                                                          │
weights update W ← W − η ∂L/∂W ◄── gradient ∂L/∂W, ∂L/∂b (chain rule, backwards) ◄─┘

Derivatives connect to machine learning with ★★★★★ fundamental strength — no gradient, no deep learning. Integrals connect with ★★★★ important strength — through probability: the true objective is an expected loss, and generative models are densities. Every arrow below is a page you can open.

Topics

Linear regression

Fit y^=w⋅x+b\hat y = w\cdot x + b by minimizing the squared error. The simplest learning problem, solvable in closed form by setting the gradient to zero — and the template for everything that follows.

UniversityApplication

Loss function

A single number L(θ)L(\theta) measuring how wrong a model with parameters θ\theta is on the data. Learning is minimizing it. Mean squared error for regression and cross-entropy for classification are the two everyone uses.

UniversityApplication

Gradient descent

Repeat θ←θ−η ∇L(θ)\theta \leftarrow \theta - \eta\,\nabla L(\theta): take a small step against the gradient. Cauchy proposed it in 1847; today it (and its stochastic, adaptive variants) trains essentially every neural network.

UniversityApplication◐ demo

Learning rate

The step size η\eta of gradient descent — the most important hyperparameter. Stability requires η<2/λmax⁡\eta < 2/\lambda_{\max} (the largest curvature); schedules (warm-up, cosine decay) start careful, go fast, and end small to settle into a minimum.

UniversityApplication

Logistic regression

Binary classification with p(y=1∣x)=σ(w⋅x+b)p(y = 1\mid x) = \sigma(w\cdot x + b), trained by maximum likelihood (cross-entropy). Convex, so gradient descent finds the global optimum — a single neuron, and the bridge to neural networks.

UniversityApplication

Activation functions

The non-linear functions applied after each linear layer. Without them a deep network collapses to one linear map. Their derivatives matter as much as their values: sigmoids saturate (vanishing gradients), ReLU passes gradients unchanged where active.

UniversityApplication

Neural networks

Compositions of affine maps and non-linearities, fθ=fL∘⋯∘f1f_\theta = f_L\circ\dots\circ f_1 with fℓ(a)=σ(Wℓa+bℓ)f_\ell(a) = \sigma(W_\ell a + b_\ell). A differentiable function with millions of adjustable parameters, trained by gradient descent.

AdvancedApplication

Automatic differentiation

Computing exact derivatives of programs by applying the chain rule to each elementary operation. Neither symbolic (no expression swell) nor numerical (no truncation error). Reverse mode gives the gradient of a scalar for a small constant times the cost of the program.

AdvancedApplication

Backpropagation

The algorithm that computes ∂L/∂W\partial L/\partial W and ∂L/∂b\partial L/\partial b for every layer of a network: one forward pass storing intermediate values, then one backward pass applying the chain rule from the loss down to the inputs. Cost: about twice the forward pass, whatever the number of parameters.

AdvancedApplication◐ demo

Stochastic gradient descent (SGD)

Use the gradient of a small random batch instead of the whole dataset: a noisy but unbiased estimate, thousands of times cheaper. The noise even helps escape saddles and sharp minima.

AdvancedApplication

Momentum and Adam

Momentum averages past gradients (a heavy ball that keeps rolling through ravines); RMSProp divides by a running RMS of the gradient (per-parameter step sizes). Adam combines both and is the default optimizer of deep learning.

AdvancedApplication

Regularization

Add a penalty to the loss, L+λ∥θ∥2L + \lambda\norm\theta^2 (weight decay) or λ∥θ∥1\lambda\norm\theta_1 (sparsity), to prefer simpler models and reduce overfitting. Its gradient 2λθ2\lambda\theta shrinks every weight a little each step.

AdvancedApplication

Loss landscape

The geometry of L(θ)L(\theta) over parameter space: valleys, plateaus, saddle points and ravines. Its curvature (Hessian) controls how fast and how stably optimizers move, and flat minima tend to generalize better than sharp ones.

AdvancedApplication

Second-order (Hessian-based) optimization

Use curvature to choose the step: θ←θ−H−1∇L\theta \leftarrow \theta - H^{-1}\nabla L. Quadratic convergence near a minimum and immune to ill-conditioning, but the Hessian of a large model cannot even be stored — so practice uses approximations (L-BFGS, Gauss–Newton, K-FAC, Shampoo).

SpecializationApplication

Deep learning

Training very deep networks (CNNs, transformers) with backpropagation and adaptive SGD on huge datasets. The calculus is the same as for one neuron; what changed is scale, architecture (residual connections, normalization, attention) and hardware.

AdvancedApplication

Convolutional networks (CNNs)

Networks whose layers convolve the input with small learned kernels: translation-equivariant, with few parameters. The backbone of computer vision since 2012.

AdvancedApplication

Support vector machines

Find the separating hyperplane with the largest margin: a convex quadratic program whose Lagrangian dual involves the data only through dot products — hence the kernel trick.

AdvancedApplication

Reinforcement learning

An agent learns to act by maximizing expected discounted reward. Value methods iterate the Bellman operator (a contraction); policy-gradient methods do gradient ascent on an expectation. RLHF uses it to fine-tune language models.

AdvancedApplication

Bayesian inference

Treat parameters as random and update a prior density into a posterior with Bayes' rule. The normalizing integral is intractable in general, so practice uses MCMC (sampling) or variational inference (optimization).

AdvancedApplication

Generative models

Models that learn a probability density of data and sample from it. Calculus is everywhere: change of variables and Jacobian determinants (flows), ELBO integrals (VAEs), Lipschitz critics (WGANs), and stochastic/ordinary differential equations integrated backwards in time (diffusion).

SpecializationApplication

Neural ODEs

A residual network xℓ+1=xℓ+hF(xℓ)x_{\ell+1} = x_\ell + hF(x_\ell) is Euler's method; let h→0h \to 0 and the network becomes an ODE x˙=Fθ(x,t)\dot x = F_\theta(x, t), evaluated by an ODE solver and trained by the adjoint method (reverse-mode AD in continuous time).

SpecializationApplication

The mathematics this domain runs on

Σ Series ★★★★★

∮ Vector calculus ★★★★★

↑ ↓ to navigate · ↵ · Esc