What is it?
The non-linear functions applied after each linear layer. Without them a deep network collapses to one linear map. Their derivatives matter as much as their values: sigmoids saturate (vanishing gradients), ReLU passes gradients unchanged where active.
Formulas
Why does it matter?
The switch from sigmoid to ReLU (around 2011) was one of the small calculus facts that made deep learning trainable.
The mathematics behind it
and : backprop reuses the forward value.
Sigmoid, softmax and tanh are all built from .
Activations must be continuous (and almost everywhere differentiable) for gradient training to work.
The perceptron's step activation is discontinuous; replacing it by the sigmoid made backpropagation possible.
ReLU has ; frameworks simply pick a value (usually 0) at the kink.
squashes to and its derivative is cheap to compute from the output.
The range of the sigmoid, , is why its output can be read as a probability.
The sigmoid saturates: as its slope tends to 0, the root of vanishing gradients.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
ℒ AI and machine learning
- Neural networks★★★★★
- Neural networks→Backpropagation★★★★★
- Neural networks→Loss landscape★★★★★
- Neural networks→Backpropagation→Deep learning★★★★★
- Neural networks→Loss landscape→Second-order (Hessian-based) optimization★★★★★
- Neural networks→Backpropagation→Deep learning→Convolutional networks (CNNs)★★★★★
- +2
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.