18.S995: Topics in Deep Learning Theory
Fall 2026 · MIT
Room change: Class will now meet in 4-163.
- Instructor
- Alex Damian
- TAs
- TBD
- Meetings
- Tu/Th 1:00pm - 2:30pm, 4-163
Course Description
This is a graduate course on the mathematical foundations of deep learning. Our goal is to develop theory that predicts what a neural network will do in an experiment we have not yet run. Topics include:
- Scaling limits: random features, the neural tangent kernel, mean-field dynamics, maximal-update parameterization, tensor programs
- Feature learning: planted models, single and multi-index models, computational-statistical gaps
- Optimization: SGD dynamics, acceleration, critical batch size, implicit regularization, the edge of stability
Learning Goals
By the end of the course, students should be able to:
- Derive the conditions for lazy training vs feature learning, relate them to different infinite-width limits, and identify the assumptions and limitations of these asymptotics
- Estimate the sample complexity of learning various features in simple planted problems, and explain how feature learning allows neural networks to outperform kernel methods
- Reason about how the hyperparameters (learning rate, batch size, momentum) affect optimization and interact with discretization and stochasticity
Prerequisites
The course assumes mathematical maturity and a working knowledge of linear algebra, analysis, and probability, including eigenvalues, matrix calculus, Gaussian random variables, concentration, and convergence arguments. Students should also be familiar with neural networks, backpropagation, and gradient descent, and comfortable training networks in JAX or PyTorch. Completion of the diagnostic Homework 0 is required to take this course.
Homework 0
Linear Algebra
Let be symmetric and positive definite, let , and define .
- Compute the gradient .
- Let follow gradient descent: . Solve for in closed form as a function of .
- For what does as ? What changes if is only positive semidefinite?
Analysis
Let be differentiable and suppose that is -Lipschitz with respect to a norm . Let satisfy , and let . Starting from , define
- Show that .
- Using , show that .
Probability
A mean-zero random variable is -sub-Gaussian if, for every ,
- Let . Show that is -sub-Gaussian.
- Let be independent mean-zero random variables such that is -sub-Gaussian. Show that is -sub-Gaussian.
- Let be -sub-Gaussian. Use Markov's inequality on and optimize over to show that, for every ,
-
Let have
i.i.d. entries. Using
Lemma 1 and a union bound, prove that, for some absolute constant
and any , with
probability at least ,
Hint: Note that for fixed
,
,
and is therefore sub-Gaussian, so you can apply part 3.
Note: This bound actually holds with , but proving this requires more advanced techniques.Lemma 1. Let and be -nets with and . Then
Proof. Let be the top left and right singular vectors of , so that , and pick and with and . Then and rearranging completes the proof.
Neural Networks
Let with . Let , , , and let be the ReLU activation, applied coordinatewise. Define the logits
-
Suppose has i.i.d.
entries, has
i.i.d. entries, and
.
- Show that each coordinate of is , and that .
- Compute .
- What is the limiting distribution of as ? Hint: CLT.
-
For a label , let
,
where is the
softmax function. Using the chain rule / backpropagation, compute
,
, and
. What are the ranks of
and
?
Hint: the gradient of is , where is the one-hot vector for . - Complete the following notebook, which involves checking your answer to part 1 and using the gradients you computed in part 2 to train the model on MNIST using gradient descent. Use only NumPy, and do not add imports beyond those provided (e.g. no JAX/PyTorch).
Assessment
There will be 3 problem sets combining theory and experiments, graded on completion. Each problem set is followed by a short in-class quiz covering the same material. The course ends with a final project. Grades are 20% problem sets, 20% quizzes, and 60% final project. Final project details will be announced during the first few weeks of the semester. There is no final exam.
AI Policy
This is a graduate topics course, and you are responsible for your own learning. AI tools are allowed on homework and the final project, with the exception of Homework 0. You are responsible for everything you submit. The homework exists to help you engage with the material. You will learn far more by working through it yourself than by handing it to a model.
Lecture notes
| Date | Topic | Notes |
|---|---|---|
| Introduction & Expressivity | ||
| Scaling Limits | ||
| Standard Parameterization | ||
| Neural Tangent Kernel | PDF Visualization | |