Power-Law Spectrum of the Random Feature Model
preprint arXiv:2603.14578 · March 2026
arXivalso in machine learning theory
I received my Ph.D. from the Mathematics department at the University of Washington (2013), under Ioana Dumitriu. After that I held an NSF postdoctoral fellowship at the Weizmann Institute of Science, in Rehovot, Israel, with Ofer Zeitouni. From 2016 to 2020 I was an assistant professor, tenure track, at Ohio State University. I joined McGill University in 2020.
My work falls into two research programs — random matrix theory and machine learning theory. The two are closely connected: the spectral theory developed for one is often exactly what the other turns out to need, and both build on a wider body of work in probability.
You might also look at some of the course notes I have given in probability, random matrix theory and machine learning, collected under notes and resources, to see the kind of work I do in these areas.
If you would like to know more about probability in Montreal:
β-ensembles, characteristic polynomials, and log-correlated fields
Random matrix theory asks what the eigenvalues and eigenvectors of a large matrix of random numbers look like. The answers are unexpectedly rigid, and the same patterns keep surfacing in unrelated places: in the sample covariance matrix a statistician forms from noisy data, in the weights a neural network is initialized with, and — first of all — in the energy levels of heavy nuclei, where Wigner replaced an intractable Hamiltonian by a random one and got the spacings right.
Much of my work concerns the $\beta$-ensembles. Dyson introduced them to sort random matrices by their symmetry, and they extend to a one-parameter family that behaves like a gas of charged particles on a line. A tridiagonal model or a canonical system turns a question about a matrix spectrum into one about a one-dimensional random operator — a close relative of the random Schrödinger operators of statistical physics.
The other thread is the characteristic polynomial. Its statistics match those of the Riemann zeta function closely enough that random matrices have become a standard source of conjectures in number theory. Its logarithm is a field that is almost — but not quite — Gaussian and log-correlated: an object with a hidden branching structure and a multifractal set of large values. That “almost” is where the difficulty and the interest both live.
Eigenvalues of a Toeplitz matrix under a vanishing random perturbation, at three noise levels.
The $\beta$-ensembles are a one-parameter family: linear statistics at $\beta = 1, 2, \pi, 10$.preprint arXiv:2603.14578 · March 2026
arXivalso in machine learning theory
preprint arXiv:2602.17960 · February 2026
arXivalso in machine learning theory
Random Matrices Theory Appl. 15(01), 2550026 · 2026
arXiv journalalso in machine learning theory
preprint arXiv:2508.20036 · August 2025
arXivalso in machine learning theory
preprint arXiv:2508.01458 · August 2025
preprint arXiv:2502.14863 · February 2025
preprint arXiv:2502.10305 · February 2025
Forum Math. Sigma 13, 126 · 2025
preprint arXiv:2403.19380 · 2024
Ann. Probab. 51(4), 1193–1248 · 2023
Analysis & PDE 16(1), 89–117 · 2023
Ann. Appl. Probab. 33(1), 549–612 · 2023
Communications in Pure and Applied Mathematics · 2022
preprint arXiv:2009.05003 · September 2020
Trans. Amer. Math. Soc. 373(7), 4999–5023 · 2020
Forum Math. Sigma 7, Paper No. e3, 72 · 2019
Electron. Commun. Probab. 23 · 2018
Probability Theory and Related Fields, 1–53 · 2018
Random Matrix Theory and Applications 7(02) · 2018
International Mathematics Research Notices 2018(16), 5028–5119 · 2018
The Annals of Probability 45(6A), 4112–4166 · 2017
Int. Math. Res. Not. IMRN, 8724–8751 · 2015
Random Matrices Theory Appl. 1(4), 60 · 2012
preprint arXiv:1707.02700 · July 2017
scaling laws, high-dimensional optimization, and the dynamics of training
Machine learning is, mathematically, a family of optimization problems that keep getting bigger, and this program is concerned with what the dynamics do in that limit. Run stochastic gradient descent on a problem whose size grows and the trajectory stops looking random: the loss curve concentrates on a deterministic path, and the effect of a learning rate, a momentum parameter or a batch size can be computed from it rather than tuned.
Scaling laws are the empirical face of the same phenomenon. They say how much better a model gets for each doubling of the resources it is given, and the observed trend is almost always a power law. What makes them interesting mathematically is that compute is no longer the only resource that matters: data is an equally important one, and the architecture decides how efficiently either is converted into performance. That is a complexity theory of a new kind.
It is also a problem about several limits at once — not just a growing dimension, but the number of parameters and the size of the dataset scaling together, with the exponent depending on how the two are taken. Two things are wanted here: a theory of stochastic optimizers sharp enough to produce those exponents rather than merely rates, and an account of the loss landscapes that give rise to power-law curves in the first place.
Loss curves for stochastic momentum on MNIST concentrating on a deterministic prediction (red) as the problem grows.
Compute-optimal scaling has phases: the exponent you observe depends on where the data and target exponents sit.
Trajectories of SGD on phase retrieval, drawn from the limiting ODE.preprint arXiv:2605.28961 · May 2026
preprint arXiv:2605.09552 · May 2026
preprint arXiv:2605.05683 · May 2026
preprint arXiv:2603.14578 · March 2026
arXivalso in random matrix theory
preprint arXiv:2602.17960 · February 2026
arXivalso in random matrix theory
preprint arXiv:2602.05298 · February 2026
Random Matrices Theory Appl. 15(01), 2550026 · 2026
arXiv journalalso in random matrix theory
preprint arXiv:2510.16687 · October 2025
preprint arXiv:2508.20036 · August 2025
arXivalso in random matrix theory
NeurIPS Spotlight · 2025
ICML · 2025
ICLR · 2025
Math. Program. 214(1-2), 1–90 · 2025
NeurIPS · 2024
NeurIPS Spotlight · 2024
Information and Inference: A Journal of the IMA 13(4), iaae028 · 2024
Harvard Data Science Review 6(3) · 2024
Electronic Communications in Probability 29, 1–15 · 2024
Advances in Neural Information Processing Systems 35 · 2022
Advances in Neural Information Processing Systems 35 · 2022
Foundations of Computational Mathematics · 2022
Advances in Neural Information Processing Systems 34, 9229–9240 · 2021
Proceedings of Thirty Fourth Conference on Learning Theory 134, 3548–3626 · 2021
SIAM Views and News 20 · 2022
random graphs, geometry, and topology
My work in probability has never been in a single program. It includes random graphs and their spectra, the topology of random complexes, point processes in hyperbolic space, and fragmentation and reinforcement.
A tree grown by preferential attachment with choice.
A random 2-dimensional cubical complex, and its fundamental group.
The Delaunay complex of a stationary point process in the hyperbolic plane.
Interval fragmentation with choice: the density of fragment sizes equidistributes.ALEA 21, 1835–1852 · 2024
Proc. Amer. Math. Soc. 151(8), 3213–3228 · 2023
J. Optim. Theory Appl. 197(1), 252–276 · 2023
The Annals of Applied Probability 32(5), 3537–3571 · 2022
Probability Theory and Related Fields · 2021
Forum of Mathematics, Sigma 9, e76 · 2021
ALEA 18, 829–854 · 2021
Combinatorics, Probability and Computing 29(2), 267–292 · 2020
International Mathematics Research Notices · 2019
Unimodularity in randomly generated graphs 719, 63–84 · 2018
Annals of Probability 46(4), 1917–1956 · 2018
Ergodic Theory and Dynamical Systems, 1–35 · 2017
Discrete & Computational Geometry 57(4), 810–823 · 2017
Israel J. Math. 212(1), 337–384 · 2016
ALEA Lat. Am. J. Probab. Math. Stat. 12(2), 903–915 · 2015
Electron. Commun. Probab. 19, no. 44, 13 · 2014
Probab. Theory Related Fields 156(3-4), 921–975 · 2013
Amer. Math. Monthly 117(2), 124–137 · 2010
Involve 2(2), 161–175 · 2009
preprint arXiv:2110.10785 · October 2021
preprint arXiv:1710.00733 · October 2017
Superseded by the published version listed above; this longer preprint remains on the arXiv.
preprint arXiv:1307.4858 · July 2013
Unpublished by intent; it remains available on the arXiv.
Course notes, exercises, and schedule: High-Dimensional Probability.
Four modules on power-law covariance, scaling laws, momentum, and spectral optimization: Course notes, slides, and resources.
Course at a glance: Syllabus.
Course at a glance: Syllabus.
Random matrix theory of high-dimensional optimization. This is a set of notes for the July 2024 random matrix theory and probability summer schools: Random matrix theory and optimization theory.
Stochastic processes notes. This is a set of notes for Math547, stochastic processes, covering Markov chains and martingales (all in discrete time). Math 547 notes.
High dimensional limits of SGD. This is a set of notes for the July 2023 summer school in probability at Lehigh university, organized by Si Tang: High-dimensional limits of SGD.