Deep learning

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

Training very deep networks (CNNs, transformers) with backpropagation and adaptive SGD on huge datasets. The calculus is the same as for one neuron; what changed is scale, architecture (residual connections, normalization, attention) and hardware.

Formulas

xℓ+1=xℓ+Fℓ(xℓ)  ⟹  ∂xℓ+1∂xℓ=I+∂Fℓ∂xℓx_{\ell+1} = x_\ell + F_\ell(x_\ell) \implies \frac{\partial x_{\ell+1}}{\partial x_\ell} = I + \frac{\partial F_\ell}{\partial x_\ell}
residual connections keep the chain of Jacobians close to the identity

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

Further reading

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc