Publications · Preprint · arXiv 2019
On the heavy-tailed theory of stochastic gradient descent for deep neural networks
arXiv preprint, 2019. Şimşekli and Gürbüzbalaban contributed equally
In brief
The long-form account of the heavy-tailed view of SGD, combining and extending the ICML 2019 tail-index measurements and the NeurIPS 2019 exit-time analysis: the gradient noise of SGD in deep networks has heavy, infinite-variance tails; the right continuous-time model is therefore a Lévy-driven stochastic differential equation rather than a diffusion; and in that model the time SGD takes to leave a basin depends on the basin’s width rather than its depth, which explains the preference for wide minima.
Topics
Cite
@misc{simsekli2019heavytailedtheory,
title = {On the heavy-tailed theory of stochastic gradient descent for deep neural networks},
author = {Umut Şimşekli and Mert Gürbüzbalaban and Thanh Huy Nguyen and Gaël Richard and Levent Sagun},
year = {2019},
howpublished = {arXiv preprint arXiv:1912.00018},
}Relatedsame topics