What is it?
Networks whose layers convolve the input with small learned kernels: translation-equivariant, with few parameters. The backbone of computer vision since 2012.
Formulas
The mathematics behind it
A conv layer slides learned kernels over the input (strictly, a cross-correlation: the kernel is not flipped).
Large convolutions can be computed by FFT; Fourier neural operators learn directly in frequency space.
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.