Deep Learning Optimization Dynamics
Understanding optimization in deep learning requires grappling with settings not captured by classical optimization theory. For example, large-batch training typically occurs in a chaotic regime called the Edge of Stability (pictured). I’ve studied how different optimizers navigate the Edge of Stability regime in order to provide simple explanations for their dynamics and behavior. You can also click this link for a fun visualization of limit cycles and chaos in Adam.