Linear regression

Level UniversityDifficulty ★★★★★Application⌖ Open in the map

What is it?

Fit y^=w⋅x+b\hat y = w\cdot x + b by minimizing the squared error. The simplest learning problem, solvable in closed form by setting the gradient to zero — and the template for everything that follows.

Formulas

min⁡w1N∥Xw−y∥2  ⟹  w∗=(X𝖳X)−1X𝖳y\min_w \frac1N\norm{Xw - y}^2 \implies w^\ast = (X^{\mathsf T}X)^{-1}X^{\mathsf T}y

Why does it matter?

Still the first model to try, the baseline for every other, and the place where the link "derivative = 0 ⇒ optimum" is clearest.

The mathematics behind it

  • Extrema in several variables★★★★★fundamental

    Ordinary least squares is solved by setting the gradient to zero: the normal equations.

Where is it used?

Computing topics reachable from here, through the chain of ideas that leads to them:

What depends on it

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc