Notes · October 2026 · 5 min read
Why heavy-tailed SGD leaves narrow basins
Gradient noise in deep learning is heavy-tailed, and that changes where SGD ends up: a Gaussian walker has to climb a barrier, a heavy-tailed one jumps it. The exit time from a basin is then set by the basin's width, not its depth — which is why heavy-tailed SGD prefers wide minima.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise (NeurIPS, 2019)
A tail-index analysis of stochastic gradient noise in deep neural networks (ICML, 2019)
The heavy-tail phenomenon in SGD (ICML, 2021)
Stochastic gradient descent does not see the loss landscape; it sees the loss landscape plus noise. The standard mental model treats that noise as Gaussian — a diffusion — and a great deal follows from the model: SGD settles into a Gibbs-like distribution, it escapes a basin at a rate that is exponentially small in the barrier height, and so, left alone, it ends up in the deepest basin it can find.
The model is wrong in an interesting way. Measured on real networks, the gradient noise of SGD has heavy tails: its tail index — the exponent in — sits well below the Gaussian value of 2, often between 1 and 1.5, and it moves with the architecture, the stepsize and the batch size. A better continuous-time model is not a diffusion but a stochastic differential equation driven by an -stable Lévy process,
whose increments have infinite variance and whose sample paths jump. Everything that follows from the Gaussian model has to be rederived, and the first thing to change is where SGD ends up.
Climbing versus jumping
Take a basin of the loss with a minimum at the bottom, a barrier of height on its rim, and a width — the distance from the minimum to the rim. For a Gaussian walker (noise scale ), Kramers' law gives the mean time to escape:
The barrier height is in the exponent, the width is nowhere. A deep basin holds a Gaussian walker essentially forever, whatever its width, and a shallow one lets it go quickly, whatever its width.
For an -stable walker the escape happens differently. It does not accumulate small steps until it has climbed the barrier; it waits for a single jump large enough to clear the rim in one go. The probability of such a jump at any given step is the tail probability of the noise beyond the distance , which is polynomial, and so is the waiting time:
Now the width is what matters and the height has disappeared — to leading order, the barrier can be as tall as you like. This is the Imkeller–Pavlyukevich asymptotics for Lévy-driven SDEs, and it is exactly what SGD inherits once its noise is heavy-tailed: the first-exit time of SGD, and of the SDE that models it, depends polynomially on the width of the basin rather than exponentially on its depth. Among several basins, the heavy-tailed walker leaves the narrow ones quickly and lingers in the wide ones; it does not care which is deepest. Since wide minima are the ones that tend to generalize, this is a mechanism — not a metaphor — connecting the tails of the gradient noise to the solutions SGD finds.
The figure shows the two laws side by side on a toy landscape. At the narrow basin is left after about a hundred steps and the wide one takes twice as long; at the narrow basin holds the walker beyond the budget of the experiment while the wide one does not. Raising the barrier of the narrow basin changes the Gaussian prediction by orders of magnitude and the heavy-tailed exit time hardly at all.
Where the tails come from
It would be a weaker story if heavy tails were only a property of some datasets. They are not. Even with Gaussian data and a quadratic loss, the iteration of SGD is a random linear recursion of Kesten type,
with a random multiplicative matrix that depends on the sampled batch. Such recursions have stationary distributions with power-law tails, and the tail index is not a property of the data: it is set by the ratio of the stepsize to the batch size, and by the stepsize schedule. SGD manufactures its own heavy tails, and the hyperparameters decide how heavy. Put the two results together and a dial appears: stepsize and batch size set , and sets which basins SGD can stay in.
A Gaussian walker climbs; a heavy-tailed one jumps. Exit times are exponential in the barrier height for the first and polynomial in the basin width for the second — and the stepsize-to-batch-size ratio decides which walker you have.
Try it yourself
The heavy-tails playground runs this experiment live: drag the tail index from 2 down to 1.1 and watch the run leave the narrow basin and settle in the wide one; raise the barrier and watch the Gaussian prediction explode while the measured heavy-tailed exit time stays put; narrow the well and watch the opposite. The readout sets the Monte Carlo exit times beside Kramers' formula and the single-jump law above.
The theorems are in the companion papers: the first-exit-time analysis of SGD under heavy-tailed noise, including the discretization error between SGD and its SDE; the measurements of the tail index of gradient noise in deep networks; and the heavy-tail phenomenon in SGD itself, with the stepsize-to-batch-size ratio controlling .
Comments and corrections are welcome by email. More notes on the notes index, or subscribe via RSS.