Theoretical Alignment Track · Course 1 of 4

Alignment Theory of Deep Learning

Why does deep learning work, and what does the answer mean for alignment? A one-week slice of the Iliad Intensive: singular learning theory, training dynamics, and data attribution.

Apply
Next cohortWeek of October 5
Intensive
1 week, Mon–Fri
Workload
25 hours, 5 per module
Format
Online, facilitated groups
Included
Lens AI Coach
Certificate
Tiered, verifiable

About this course

Deep learning works far better than classical theory says it should. Networks with more parameters than data points still generalize, and gradient descent still finds good solutions on non-convex losses. This course studies the mathematics that tries to explain why, and asks what those explanations mean for aligning the systems we train.

The week starts with the alignment problem itself, so that the theory has a purpose. It then moves through the open mysteries of deep learning, singular learning theory, the exact training dynamics of deep linear networks, and data attribution. Each unit uses the Iliad Intensive lecture notes and worksheets, followed by a small-group session.

This is the theory track, not the engineering track. You will do derivations and exercises by hand more often than you will write code. You need solid linear algebra, multivariable calculus, and probability, and you should already know how neural networks are trained.

Who is it for:

Mathematicians and physicists

You have a strong quantitative background and want to see where it applies in AI alignment research.Apply your maths to alignment research.

Theoretical computer scientists

You want a rigorous account of learning, generalization, and inductive bias, beyond what empirical ML papers give you.A rigorous account of why networks learn.

ML engineers who want the theory

You train models already and want to understand loss landscapes, implicit bias, and why your training runs behave as they do.Understand why your training runs behave as they do.

Aspiring alignment researchers

You are considering agendas such as developmental interpretability or singular learning theory and want the foundations before you apply to a fellowship.Build foundations before a research fellowship.

Curriculum

5 UNITS · ~5H EACH

A one-week slice of the month-long Iliad Intensive: its alignment introduction, then its full cluster on the theory of learning. The selection follows the recommendation of the Iliad curriculum lead.

01

AI alignment introduction

Choose what an AI system should be aligned to, from corrigibility to intent alignment. Decompose the problem into outer and inner misalignment with training stories. Examine goal-directedness and instrumental convergence, then survey the range of views on how hard the problem is.

02

Mysteries of deep learning

Three classical puzzles: why networks approximate efficiently, why they generalize despite overparameterization, and why SGD succeeds on non-convex losses. Add two empirical ones, shared representations across architectures and in-context learning, and evaluate the candidate explanations.

03

Singular learning theory

Treat degeneracy as the core of how neural networks learn. Work from the parameter–function map and continuous symmetries to the local learning coefficient, then to Watanabe’s free energy formula and Bayesian phase transitions.

04

Training dynamics

Solve the learning dynamics of deep linear networks exactly. Cover loss-landscape geometry, conserved quantities under gradient flow, and the neural tangent kernel. Contrast the rich saddle-to-saddle regime with the lazy regime, and see how their mixture explains grokking.

05

Data attribution

Ask which training examples caused a model’s behavior. Start from counterfactuals and Shapley values, derive classical influence functions and the fixes that make them work for neural networks, then study Bayesian influence functions and training-dynamics unrolling.

Course Outcomes

A clear statement of the alignment problem

You can explain alignment targets, outer and inner misalignment, and why goal-directed systems raise additional risks.

A map of what theory cannot yet explain

You can state the open mysteries of approximation, generalization, and optimization, and judge proposed explanations for each.

Working knowledge of singular learning theory

You can compute degeneracy in small models and explain what the local learning coefficient measures and why it matters.

Exact models of training

You can derive how deep linear networks learn, and use the rich and lazy regimes to reason about feature learning in real networks.

Tools to trace behavior to data

You can derive influence functions, state when their approximations fail, and compare them with Bayesian and unrolling methods.

A path into theory-driven alignment research

You know the core papers and open questions well enough to follow current work and choose what to study next.

Upcoming cohorts

INTENSIVE

5h/day · 1 week

Week of October 5Applications open here soon.

Common questions

Not sure it’s for you? Come to the first session. If it’s not the right fit, we’ll help you find something that is.

Don’t wait 3 months for a 5% acceptance rate.

Upskill now

Next cohort · week of Oct 5

Apply