Alex Damian

I am an Assistant Professor at MIT with a shared appointment between Mathematics and EECS[AI+D], and a member of LIDS. Before joining MIT, I was a Kempner Research Fellow at the Kempner Institute at Harvard University.

I received my Ph.D. in Applied and Computational Mathematics at Princeton University under the supervision of Jason D. Lee, where I was supported by a Jane Street Graduate Research Fellowship and an NSF Graduate Research Fellowship. Before that, I received my B.S. in Mathematics at Duke University, where I was fortunate to work with Cynthia Rudin and Hau-Tieng Wu.

Research

My research is focused on developing a predictive, and ultimately prescriptive, theory of deep learning. Selected papers are listed under each research direction below. A full list is on Google Scholar.

Deep Learning Optimization Dynamics

Understanding optimization in deep learning requires grappling with settings not captured by classical optimization theory. For example, large-batch training typically occurs in a chaotic regime called the Edge of Stability (pictured). I’ve studied how different optimizers navigate the Edge of Stability regime in order to provide simple explanations for their dynamics and behavior. You can also click this link for a fun visualization of limit cycles and chaos in Adam.

Edge of Stability

Representation Learning in Simple Models

The miracle of deep learning is that neural networks automatically extract meaningful representations from raw data during optimization. To gain insights into this process, I’ve studied the optimization dynamics of simple models trained on synthetic data to ask: What representations are learned? How many samples does the network need to learn them? What signals in the gradient help guide optimization towards them? I’ve worked on these questions in both feed-forward neural networks (MLPs) and Transformers.

Induction Heads

Computational-to-Statistical Gaps

Many high-dimensional learning problems exhibit a conjectured gap between the minimum number of samples needed information-theoretically to solve the problem, and the number of samples needed by polynomial time algorithms. This implies a fundamental tradeoff between runtime and sample complexity. I’ve studied this tradeoff in Gaussian single-index and multi-index models to identify structures that can make learning problems hard or easy.

Landscape Smoothing

Recruiting

I am actively recruiting Ph.D. students to start in Fall 2027. Interested students should apply to the Mathematics and/or EECS departments at MIT and list my name in their application.

I am not currently hiring postdocs or supervising students outside the Ph.D. program.