Mert Gürbüzbalaban

Rutgers Business School · Department of Management Science and Information Systems

Mert Gürbüzbalaban


Associate Professor at Rutgers University. I design and analyze optimization algorithms for large-scale optimization problems and machine learning — fast first-order methods that stay robust to noisy, biased, or adversarial gradients; the stochastic dynamics behind heavy tails and Langevin sampling; minimax optimization and primal–dual methods; and learning over networks — with tools from convex optimization, probability, and control.

Graduate faculty in Statistics and in Electrical & Computer Engineering · Member of DIMACS

66 publications · ICML 2019 Best Paper Honorable Mention · Research supported by the NSF and the ONR

Portrait of Mert Gürbüzbalaban
Researchfour directions

Current research directions

One question organizes each direction; together they form a single program — optimization algorithms treated as dynamical systems driven by imperfect, stochastic, or adversarial information.

Direction I

Robust and risk-sensitive optimization

How should fast algorithms be designed when gradients are noisy, biased, or adversarial?

Direction II

Stochastic dynamics and sampling

Which stochastic dynamics make optimization and sampling faster, more stable — or heavy-tailed?

Direction III

Minimax optimization and primal–dual methods

How should saddle-point problems — the form that robust learning, adversarial training, and constrained problems take — be solved fast and reliably from stochastic gradients?

Direction IV

Distributed learning and large-scale computation

How do communication, computation, robustness, and stochasticity interact over networks?

Optimization algorithmsRandomness in the gradientAdversaries and biased estimatesHeavy tails · Langevin samplingReshuffling and incremental methodsRobust acceleration · riskMinimax and primal–dual methodsDynamical systems + probabilityDistributed computation over networks
How the areas belong to one program: randomness and adversaries in the gradient both enter a dynamical-systems-and-probability toolbox, which then scales to networks. Boxes link to the research page.
Featured program2019 – present

One dynamical system, three norms

A running theme of my current work: an optimization algorithm is a feedback dynamical system, and once you take that literally, its behavior under imperfect gradients becomes something you can compute — not just bound. Average-case fragility is an H2 norm. Worst-case amplification of adversarial errors is an H∞ norm, with an explicit worst-case noise achieving it. Rare-event behavior is a risk-sensitive index whose convex conjugate is a large-deviations rate function. Each is computable for the generalized momentum family, and each can be designed for.

Lens I — Random, unbiased noise

H2 norm + risk measures

Honest but noisy gradients. Iterates equilibrate to a stationary distribution; its size solves a Lyapunov equation, and entropic risk measures capture its tails.

Lens II — Adversarial, deterministic noise

H∞ norm

Worst-case errors with finite energy. The H∞ norm of the algorithm is computable in closed form on quadratics, together with the worst-case noise itself.

Lens III — Biased, random noise

Risk-sensitive index + LDP

Systematically biased gradient estimates. A Riccati reduction yields exact risk-sensitive indices, and large deviations of the running suboptimality follow by convex conjugacy.

Speed–robustness–risk trade-offs in first-order methods are fundamental — and designable.

Try the interactive momentum playground →

The full research program →

Selected papersof 66 works

Selected publications

J. Nonlinear Var. Anal.2026

Accelerated gradient methods with biased gradient estimates: risk sensitivity, high-probability guarantees, and large deviation bounds

Mert Gürbüzbalaban, Yasa Syed, Necdet Serhat Aybat

Journal of Nonlinear and Variational Analysis (special issue), 2026.

Studies accelerated methods when gradient estimates are biased as well as noisy. Computes the risk-sensitive index of generalized momentum methods via a Riccati-equation reduction, characterizes when it blows up, and derives a large deviation principle whose rate function is the convex conjugate of that index — turning rare-event behavior of the running suboptimality into a designable quantity, with finite-time high-probability guarantees beyond quadratics.

@article{gurbuzbalaban2026biased,
  title   = {Accelerated gradient methods with biased gradient estimates: risk sensitivity, high-probability guarantees, and large deviation bounds},
  author  = {Mert Gürbüzbalaban and Yasa Syed and Necdet Serhat Aybat},
  year    = {2026},
  journal = {Journal of Nonlinear and Variational Analysis (special issue)},
  note    = {arXiv:2509.13628},
}
Optim. Methods Softw.2025

Entropic risk-averse generalized momentum methods

Bugra Can, Mert Gürbüzbalaban

Optimization Methods and Software, 40(6), pp. 1535–1583, 2025.

Builds a unified convergence and risk analysis for the generalized momentum family (covering gradient descent, heavy ball, and Nesterov acceleration) under stochastic gradient errors, bounding the entropic risk and entropic value-at-risk of suboptimality. The bounds support risk-averse parameter selection (RA-GMM): choosing step and momentum on the rate–risk Pareto frontier rather than for expected performance alone.

@article{can2025entropic,
  title   = {Entropic risk-averse generalized momentum methods},
  author  = {Bugra Can and Mert Gürbüzbalaban},
  year    = {2025},
  journal = {Optimization Methods and Software},
  volume  = {40(6)},
  pages   = {1535--1583},
  doi     = {10.1080/10556788.2025.2549356},
}
Math. Oper. Res.2025

Robustly stable accelerated momentum methods with a near-optimal L2 gain and H∞ performance

Mert Gürbüzbalaban

Mathematics of Operations Research, 2025. Dedicated to Michael L. Overton.

Treats first-order algorithms as feedback dynamical systems subject to worst-case, finite-energy gradient errors and computes their H∞ norm in closed form on quadratics — together with the explicit worst-case noise achieving it (a decaying cosine). Introduces robustly stable variants (RS-GD, RS-HB) with near-optimal L2 gain, quantifies real stability radii of momentum methods, and extends the guarantees beyond quadratics through matrix-inequality certificates.

@article{gurbuzbalaban2025hinf,
  title   = {Robustly stable accelerated momentum methods with a near-optimal {$L_2$} gain and {$H_\infty$} performance},
  author  = {Mert Gürbüzbalaban},
  year    = {2025},
  journal = {Mathematics of Operations Research},
  doi     = {10.1287/moor.2023.0321},
}
Oper. Res.2022

Provides non-asymptotic global convergence guarantees for stochastic gradient Hamiltonian Monte Carlo on non-convex problems, quantifying when and how momentum accelerates Langevin-based optimization.

@article{gao2022sghmc,
  title   = {Global convergence of stochastic gradient Hamiltonian Monte Carlo for nonconvex stochastic optimization: nonasymptotic performance bounds and momentum-based acceleration},
  author  = {Xuefeng Gao and Mert Gürbüzbalaban and Lingjiong Zhu},
  year    = {2022},
  journal = {Operations Research},
  volume  = {70(5)},
  pages   = {2931--2947},
  doi     = {10.1287/opre.2021.2162},
}
J. Mach. Learn. Res.2021

Decentralized stochastic gradient Langevin dynamics and Hamiltonian Monte Carlo

Mert Gürbüzbalaban, Xuefeng Gao, Yuanhan Hu, Lingjiong Zhu

Journal of Machine Learning Research, 22(239), pp. 1–69, 2021.

Introduces decentralized versions of stochastic gradient Langevin dynamics and Hamiltonian Monte Carlo for Bayesian learning over networks of agents, with non-asymptotic guarantees on sampling accuracy.

@article{gurbuzbalaban2021decentralizedlangevin,
  title   = {Decentralized stochastic gradient Langevin dynamics and Hamiltonian Monte Carlo},
  author  = {Mert Gürbüzbalaban and Xuefeng Gao and Yuanhan Hu and Lingjiong Zhu},
  year    = {2021},
  journal = {Journal of Machine Learning Research},
  volume  = {22(239)},
  pages   = {1--69},
  note    = {arXiv:2007.00590},
}
ICML2021

The heavy-tail phenomenon in SGD

Mert Gürbüzbalaban, Umut Şimşekli, Lingjiong Zhu

International Conference on Machine Learning (ICML), pp. PMLR 139:3964–3975, 2021.

Shows that SGD can produce heavy-tailed, power-law iterate fluctuations even from light-tailed data: multiplicative gradient noise drives a Kesten-type random recursion whose stationary law has a tail index controlled by the stepsize-to-batch-size ratio — connecting algorithm hyperparameters to the tail behavior, and thereby the generalization properties, of the learned solutions.

@inproceedings{gurbuzbalaban2021heavytail,
  title   = {The heavy-tail phenomenon in SGD},
  author  = {Mert Gürbüzbalaban and Umut Şimşekli and Lingjiong Zhu},
  year    = {2021},
  booktitle = {International Conference on Machine Learning (ICML)},
  pages   = {PMLR 139:3964--3975},
  note    = {arXiv:2006.04740},
}
Group3 current · 6 alumni

Research group

Current Ph.D. students: Hamza Rafi, Meija Chen, Mustafa Ali Kutbay. Graduates have gone on to Amazon, Wells Fargo, Barclays, UC San Diego, and Université Côte d’Azur. Students and alumni →

Prospective students

I work with Ph.D. students on optimization, stochastic algorithms, machine-learning theory, sampling, and related areas. Students interested in these topics are welcome to contact me with a short description of their background and research interests. mg1366@rutgers.edu

Notescompanions to the papers

Latest research note

July 2026

How fragile is acceleration?

Momentum buys a quadratic speedup — and pays for it in sensitivity to gradient errors. A guided tour of why that trade-off is a theorem, not an engineering accident, and how to design on the frontier instead of falling off it.

News

Recent news

Oct 2026
Upcoming

Invited talk, “Robust and Risk-Sensitive Acceleration in Gradient Methods,” at the Financial/Actuarial Mathematics Seminar, University of Michigan, Ann Arbor (October 7). Details →

Sep 2026

“RESIST: resilient decentralized learning using consensus gradient descent” (with C. Fang, R. Dixit, and W. U. Bajwa) appears in Transactions on Machine Learning Research with a Featured certification. →

2026

“Accelerated gradient methods with biased gradient estimates: risk sensitivity, high-probability guarantees, and large deviation bounds” (with Y. Syed and N. S. Aybat) appears in the Journal of Nonlinear and Variational Analysis. →

2026

“Mean-semideviation-based distributionally robust learning with weakly convex losses” (with L. Zhu and A. Ruszczyński) appears in Mathematical Programming 215(1). →

2025

“Robustly stable accelerated momentum methods with a near-optimal L2 gain and H∞ performance” appears in Mathematics of Operations Research. The paper is dedicated to Michael L. Overton. →