What is it?
An agent learns to act by maximizing expected discounted reward. Value methods iterate the Bellman operator (a contraction); policy-gradient methods do gradient ascent on an expectation. RLHF uses it to fine-tune language models.
Formulas
- Bellman optimality equation
- policy gradient theorem
The mathematics behind it
Discounting with makes the infinite return a convergent geometric-type series.
Agents maximize expected discounted return; policy gradients differentiate an expectation.
Value iteration converges because the Bellman operator is a -contraction.
This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.