What is it?
Training very deep networks (CNNs, transformers) with backpropagation and adaptive SGD on huge datasets. The calculus is the same as for one neuron; what changed is scale, architecture (residual connections, normalization, attention) and hardware.
Formulas
- residual connections keep the chain of Jacobians close to the identity
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
What depends on it
Further reading
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.