What is it?
Fit by minimizing the squared error. The simplest learning problem, solvable in closed form by setting the gradient to zero — and the template for everything that follows.
Formulas
Why does it matter?
Still the first model to try, the baseline for every other, and the place where the link "derivative = 0 ⇒ optimum" is clearest.
The mathematics behind it
Ordinary least squares is solved by setting the gradient to zero: the normal equations.
Where is it used?
Computing topics reachable from here, through the chain of ideas that leads to them:
ℒ AI and machine learning
- Logistic regression★★★★★
- Logistic regression→Neural networks★★★★★
- Logistic regression→Neural networks→Backpropagation★★★★★
- Logistic regression→Neural networks→Loss landscape★★★★★
- Logistic regression→Neural networks→Backpropagation→Deep learning★★★★★
- Logistic regression→Neural networks→Loss landscape→Second-order (Hessian-based) optimization★★★★★
- +3
What depends on it
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.