AI and machine learning
The calculus behind modern AI: loss functions, gradients, backpropagation, optimizers and probabilistic models.
21 topics
The calculus behind modern AI
Strip a neural network of its jargon and what remains is calculus: a big differentiable function, a scalar loss that measures its error, and an algorithm that follows the derivative of the loss downhill.
Calculus
├── Derivatives ──► Gradients ──► Optimization ──► Gradient descent
├── Partial derivatives + chain rule ──► Backpropagation
├── Integrals ──► Probability ──► Probabilistic ML (likelihoods, Bayes, diffusion)
└── Linear algebra + calculus ──► Neural networks
One training step, end to end:
input x ─► layer z = Wx + b ─► activation a = σ(z) ─► prediction ŷ ─► loss L(ŷ, y)
│
weights update W ← W − η ∂L/∂W ◄── gradient ∂L/∂W, ∂L/∂b (chain rule, backwards) ◄─┘
Derivatives connect to machine learning with ★★★★★ fundamental strength — no gradient, no deep learning. Integrals connect with ★★★★ important strength — through probability: the true objective is an expected loss, and generative models are densities. Every arrow below is a page you can open.
Topics
Linear regression
Fit by minimizing the squared error. The simplest learning problem, solvable in closed form by setting the gradient to zero — and the template for everything that follows.
Loss function
A single number measuring how wrong a model with parameters is on the data. Learning is minimizing it. Mean squared error for regression and cross-entropy for classification are the two everyone uses.
Gradient descent
Repeat : take a small step against the gradient. Cauchy proposed it in 1847; today it (and its stochastic, adaptive variants) trains essentially every neural network.
Learning rate
The step size of gradient descent — the most important hyperparameter. Stability requires (the largest curvature); schedules (warm-up, cosine decay) start careful, go fast, and end small to settle into a minimum.
Logistic regression
Binary classification with , trained by maximum likelihood (cross-entropy). Convex, so gradient descent finds the global optimum — a single neuron, and the bridge to neural networks.
Activation functions
The non-linear functions applied after each linear layer. Without them a deep network collapses to one linear map. Their derivatives matter as much as their values: sigmoids saturate (vanishing gradients), ReLU passes gradients unchanged where active.
Neural networks
Compositions of affine maps and non-linearities, with . A differentiable function with millions of adjustable parameters, trained by gradient descent.
Automatic differentiation
Computing exact derivatives of programs by applying the chain rule to each elementary operation. Neither symbolic (no expression swell) nor numerical (no truncation error). Reverse mode gives the gradient of a scalar for a small constant times the cost of the program.
Backpropagation
The algorithm that computes and for every layer of a network: one forward pass storing intermediate values, then one backward pass applying the chain rule from the loss down to the inputs. Cost: about twice the forward pass, whatever the number of parameters.
Stochastic gradient descent (SGD)
Use the gradient of a small random batch instead of the whole dataset: a noisy but unbiased estimate, thousands of times cheaper. The noise even helps escape saddles and sharp minima.
Momentum and Adam
Momentum averages past gradients (a heavy ball that keeps rolling through ravines); RMSProp divides by a running RMS of the gradient (per-parameter step sizes). Adam combines both and is the default optimizer of deep learning.
Regularization
Add a penalty to the loss, (weight decay) or (sparsity), to prefer simpler models and reduce overfitting. Its gradient shrinks every weight a little each step.
Loss landscape
The geometry of over parameter space: valleys, plateaus, saddle points and ravines. Its curvature (Hessian) controls how fast and how stably optimizers move, and flat minima tend to generalize better than sharp ones.
Second-order (Hessian-based) optimization
Use curvature to choose the step: . Quadratic convergence near a minimum and immune to ill-conditioning, but the Hessian of a large model cannot even be stored — so practice uses approximations (L-BFGS, Gauss–Newton, K-FAC, Shampoo).
Deep learning
Training very deep networks (CNNs, transformers) with backpropagation and adaptive SGD on huge datasets. The calculus is the same as for one neuron; what changed is scale, architecture (residual connections, normalization, attention) and hardware.
Convolutional networks (CNNs)
Networks whose layers convolve the input with small learned kernels: translation-equivariant, with few parameters. The backbone of computer vision since 2012.
Support vector machines
Find the separating hyperplane with the largest margin: a convex quadratic program whose Lagrangian dual involves the data only through dot products — hence the kernel trick.
Reinforcement learning
An agent learns to act by maximizing expected discounted reward. Value methods iterate the Bellman operator (a contraction); policy-gradient methods do gradient ascent on an expectation. RLHF uses it to fine-tune language models.
Bayesian inference
Treat parameters as random and update a prior density into a posterior with Bayes' rule. The normalizing integral is intractable in general, so practice uses MCMC (sampling) or variational inference (optimization).
Generative models
Models that learn a probability density of data and sample from it. Calculus is everywhere: change of variables and Jacobian determinants (flows), ELBO integrals (VAEs), Lipschitz critics (WGANs), and stochastic/ordinary differential equations integrated backwards in time (diffusion).
Neural ODEs
A residual network is Euler's method; let and the network becomes an ODE , evaluated by an ODE solver and trained by the adjoint method (reverse-mode AD in continuous time).
The mathematics this domain runs on
ƒ Foundations ★★★★★
- Functions★★★★★→Neural networks
- Composition★★★★★→Neural networks, Automatic differentiation
- Logarithmic functions★★★★★→Loss function
- Inverse functions★★★★★→Generative models
- Exponential functions★★★★★→Activation functions
- Hyperbolic functions★★★★★→Activation functions
- Absolute value★★★★★→Loss function
- Domain and range★★★★★→Activation functions
ε Continuity ★★★★★
f′ Derivatives ★★★★★
- Derivative★★★★★→Gradient descent, Automatic differentiation
- Differentiation rules★★★★★→Automatic differentiation
- Derivatives of elementary functions★★★★★→Activation functions, Automatic differentiation
- Chain rule★★★★★→Automatic differentiation, Backpropagation
- Differentiability and one-sided derivatives★★★★★→Loss function, Activation functions
- Higher-order derivatives★★★★★→Second-order (Hessian-based) optimization
min Extrema and optimization ★★★★★
- Maxima and minima★★★★★→Loss function
- Convexity and concavity★★★★★→Loss function, Loss landscape, Support vector machines
- Lagrange multipliers★★★★★→Regularization, Support vector machines
- KKT conditions★★★★★→Support vector machines
- Critical points and derivative tests★★★★★→Gradient descent, Loss landscape
≈ Numerical methods ★★★★★
- Newton's method★★★★★→Second-order (Hessian-based) optimization
- Conditioning★★★★★→Loss landscape
- Numerical stability★★★★★→Loss function
- Order of convergence★★★★★→Gradient descent
- Fixed-point iteration★★★★★→Reinforcement learning
- Numerical differentiation★★★★★→Automatic differentiation
- Numerical integration (quadrature)★★★★★→Bayesian inference
Tₙ Taylor series ★★★★★
Σ Series ★★★★★
∇ Multivariable calculus ★★★★★
- Functions of several variables★★★★★→Loss function
- Partial derivatives★★★★★→Backpropagation
- Gradient★★★★★→Gradient descent
- Multivariable chain rule★★★★★→Automatic differentiation, Backpropagation
- Jacobian matrix★★★★★→Automatic differentiation, Generative models
- Hessian matrix★★★★★→Loss landscape, Second-order (Hessian-based) optimization
- Extrema in several variables★★★★★→Linear regression, Loss landscape
- Surfaces and level sets★★★★★→Loss landscape
- +2
ẏ Differential equations ★★★★★
ℱ Transforms ★★★★★
ℙ Probability and calculus ★★★★★
- Probability density function★★★★★→Loss function, Bayesian inference, Generative models
- Expectation★★★★★→Loss function, Stochastic gradient descent (SGD), Reinforcement learning
- Maximum likelihood estimation★★★★★→Loss function, Logistic regression, Generative models
- Continuous random variables★★★★★→Generative models
- Variance★★★★★→Neural networks, Stochastic gradient descent (SGD), Momentum and Adam
- Continuous distributions★★★★★→Generative models