Mert Gürbüzbalaban

Publications · Preprint · arXiv 2019

On the heavy-tailed theory of stochastic gradient descent for deep neural networks


Umut Şimşekli, Mert Gürbüzbalaban, Thanh Huy Nguyen, Gaël Richard, Levent Sagun

arXiv preprint, 2019. Şimşekli and Gürbüzbalaban contributed equally

In brief

The long-form account of the heavy-tailed view of SGD, combining and extending the ICML 2019 tail-index measurements and the NeurIPS 2019 exit-time analysis: the gradient noise of SGD in deep networks has heavy, infinite-variance tails; the right continuous-time model is therefore a Lévy-driven stochastic differential equation rather than a diffusion; and in that model the time SGD takes to leave a basin depends on the basin’s width rather than its depth, which explains the preference for wide minima.

Cite
@misc{simsekli2019heavytailedtheory,
  title   = {On the heavy-tailed theory of stochastic gradient descent for deep neural networks},
  author  = {Umut Şimşekli and Mert Gürbüzbalaban and Thanh Huy Nguyen and Gaël Richard and Levent Sagun},
  year    = {2019},
  howpublished = {arXiv preprint arXiv:1912.00018},
}

← All publications · Research program · Mert Gürbüzbalaban