Reinforcement learning

Level AdvancedDifficulty ★★★★★Application⌖ Open in the map

What is it?

An agent learns to act by maximizing expected discounted reward. Value methods iterate the Bellman operator (a contraction); policy-gradient methods do gradient ascent on an expectation. RLHF uses it to fine-tune language models.

Formulas

V(s)=max⁡a(r(s,a)+γ∑s′P(s′∣s,a) V(s′))V(s) = \max_a\Big(r(s,a) + \gamma\sum_{s'}P(s'\mid s,a)\,V(s')\Big)
Bellman optimality equation
∇θJ=𝔼πθ[∇θlog⁡πθ(a∣s) Gt]\nabla_\theta J = \E_{\pi_\theta}\big[\nabla_\theta\log\pi_\theta(a\mid s)\,G_t\big]
policy gradient theorem

The mathematics behind it

  • Geometric series★★★★★fundamental

    Discounting with γ<1\gamma < 1 makes the infinite return a convergent geometric-type series.

  • Expectation★★★★★fundamental

    Agents maximize expected discounted return; policy gradients differentiate an expectation.

  • Fixed-point iteration★★★★★fundamental

    Value iteration converges because the Bellman operator is a γ\gamma-contraction.

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc