Neural ODEs

Level SpecializationDifficulty ★★★★★Application⌖ Open in the map

What is it?

A residual network xℓ+1=xℓ+hF(xℓ)x_{\ell+1} = x_\ell + hF(x_\ell) is Euler's method; let h→0h \to 0 and the network becomes an ODE x˙=Fθ(x,t)\dot x = F_\theta(x, t), evaluated by an ODE solver and trained by the adjoint method (reverse-mode AD in continuous time).

Formulas

dhdt=fθ(h,t),dadt=−a𝖳∂fθ∂h,a=∂L∂h\frac{\dd h}{\dd t} = f_\theta(h, t), \qquad \frac{\dd a}{\dd t} = -a^{\mathsf T}\frac{\partial f_\theta}{\partial h}, \quad a = \frac{\partial L}{\partial h}

The mathematics behind it

  • Ordinary differential equations★★★★★fundamental

    A neural ODE defines the hidden state by h′(t)=fθ(h,t)h'(t) = f_\theta(h, t) and calls an ODE solver as a layer.

  • Runge–Kutta methods★★★★★frequent

    Neural ODE libraries (torchdiffeq, Diffrax) default to adaptive Runge–Kutta solvers.

This page has the essentials. A fuller treatment (intuition, formal definition, worked example) is on the way.

↑ ↓ to navigate · ↵ · Esc