Interactive
Heavy tails and basin exits
Gradient noise in deep learning is heavy-tailed, and that changes where SGD ends up. This page runs SGD on a landscape with a narrow, deep basin and a wide, shallow one, under α-stable noise whose tail index you set, and measures how long each basin holds the iterate.
Start with Gaussian noise. The run sits in the narrow basin for all 6,000 steps, and in the exit experiment most runs never escape within their budget: a Gaussian walker has to climb the barrier, and the waiting time for that grows like exp(Δ/ησ²) in the barrier height Δ (Kramers’ law). Now lower the tail index α. The trajectory develops jumps, and the first one that clears the barrier is enough — the mean exit time from the narrow basin drops by orders of magnitude, and the run settles in the wide basin instead.
Then press Taller barrier. The Gaussian prediction explodes, while the heavy-tailed exit times hardly move: a jump does not care how tall the barrier is, only how far away it is. Press Narrower well and the opposite happens — the heavy-tailed exit time falls with the width of the basin, polynomially, while the Gaussian one barely notices. Width, not depth, is what heavy-tailed SGD sees.
Interactive · heavy tails and basin exits
- Exit time, narrow well
- …
- mean over 24 runs at α = 1.5
- Exit time, wide well
- …
- same noise, started at the wide minimum
- Gaussian prediction
- 4.8e4
- Kramers, narrow well: ∝ exp(Δ/ησ²) — exponential in the barrier height Δ = 1.17
- Heavy-tail prediction
- 316
- one jump, narrow well: ∝ (a/ησ)α — polynomial in the distance a = 0.95 to the barrier
The run is SGD, xk+1 = xk − η (f′(xk) + σ ξk), with η = 0.02 and ξk symmetric α-stable of unit scale — Gaussian at α = 2, Cauchy at α = 1 — drawn with the Chambers–Mallows–Stuck sampler. The landscape is a wide bowl with a narrow dimple on its slope; the sliders set the dimple’s width and depth. Exit times are measured from the minimum until the iterate first crosses the barrier; each run has a budget of 8,000 steps, so hollow markers are lower bounds. The two predictions are Kramers’ formula for Gaussian noise and the single-jump asymptotics for α < 2, both for the narrow well. Lower α, then raise the barrier: the Gaussian prediction explodes while the heavy-tailed exits hardly change — the iterate does not climb the barrier, it jumps it.
A Gaussian walker climbs; a heavy-tailed one jumps. Exit times are exponential in the barrier height for the first and polynomial in the basin width for the second — which is why heavy-tailed SGD ends up in wide basins.
The theory behind this page is in NeurIPS 2019, ICML 2019, and ICML 2021: the first analyzes the first-exit times of SGD modeled as a Lévy-driven stochastic differential equation, building on the Imkeller–Pavlyukevich asymptotics for α-stable noise, and shows that the exit time depends polynomially on the width of the basin rather than exponentially on its depth; the second measures the tail index of gradient noise in deep networks and finds it well below 2; the third shows that SGD can generate heavy tails on its own, with the stepsize-to-batch-size ratio setting α. The wider thread is on the research page.
To reference this page:
@misc{gurbuzbalaban2026heavytails,
title = {Heavy tails and basin exits, interactively},
author = {G{\"u}rb{\"u}zbalaban, Mert},
year = {2026},
howpublished = {\url{https://mert-g.org/playground/heavy-tails/}}
}