What is it?
Find the separating hyperplane with the largest margin: a convex quadratic program whose Lagrangian dual involves the data only through dot products — hence the kernel trick.
Formulas
- the dual: only kernels appear
The mathematics behind it
Training an SVM is a convex quadratic program: a unique global optimum.
The SVM dual is obtained with Lagrange multipliers; only points with (support vectors) matter.
Complementary slackness is why only the support vectors (points on or inside the margin) get non-zero weight.
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.