ML engineers
You want to build the evaluations, monitors, and training setups that catch collusion, sandbagging, and sabotage.Build the evals and monitors that catch hidden failures.
AI Control Track · Course 2 of 2
Collusion, sandbagging, research sabotage, and threat models beyond scheming. The advanced second half of the AI control track.
AI Control 1 made the case for control and built its high-stakes toolkit: monitoring, resampling, and the settings where protocols are tested. AI Control 2 starts there and asks what happens when failure is harder to see: a monitor that colludes with the model it watches, a model that underperforms on purpose, and sabotage spread across many ordinary tasks.
Across five units, you’ll work through guided exercises on collusion and legibility, stress-test capability elicitation, study a benchmark for research sabotage in ML code, and measure reward seeking. Then you’ll look past scheming altogether: reward seekers, spillway motivations, and whether a deal with an early misaligned AI could work.
This course assumes AI Control 1 and starts from its results. The last unit shows where to go next in control research: fellowships, organizations, and open roles.
You want to build the evaluations, monitors, and training setups that catch collusion, sandbagging, and sabotage.Build the evals and monitors that catch hidden failures.
You finished AI Control 1 and want to reach the open questions: low-stakes control, exploration hacking, and threat models beyond scheming.Reach the open questions in control research.
You test models for dangerous capabilities and want to know when a model hides what it can do, and how to draw it out.Find out when a model hides what it can do.
You want to judge which threat models matter, what control can and cannot cover, and whether deals with misaligned AIs belong in a plan.Judge threat models, limits, and deals with AIs.
Adapted from the second half of the AI Safety Tracks Control curriculum. Each unit combines primary readings and guided exercises with a group discussion. Optional material lets you go further.
Take the red team's seat in two guided exercises: collusion between an untrusted model and an untrusted monitor, and legibility protocols that make work easier for a trusted monitor to check. Then shift to low-stakes control, where failures only matter in aggregate.
Study strategic underperformance as a threat model: sandbagging, and exploration hacking during RL training. Stress-test capability elicitation in a guided exercise, and read the paper on whether models can learn to resist RL training.
Work with a benchmark for detecting and fixing sabotage in ML research codebases, then complete the long guided exercise on exploration hacking. An optional reading covers the full capability elicitation paper.
Consider threat models that are not schemers: reward seekers, cheaply satisfied preferences, and channelling reward hacking into a spillway motivation. Review the empirical evidence, then measure reward seeking in a guided exercise.
Ask why you might trade with early misaligned AIs, what a deal with a schemer could look like, and what stands in the way. Finish with a map of research programs, organizations, and opportunity boards for continuing in AI control.
Explain how an untrusted model and its monitor could collude, and what makes a model's output legible enough for a trusted monitor to check.
Tell high-stakes from low-stakes control, and explain which defences still work when harm only shows up in aggregate.
Recognise sandbagging and exploration hacking, and judge whether an evaluation draws out what a model can really do.
Know how sabotage in ML research code is benchmarked, detected, and fixed, and where current methods fall short.
Compare schemers with reward seekers and other motivations, and say what each one means for the design of a control protocol.
Weigh the case for deals with early misaligned AIs, and leave with a map of programs and organizations to continue with.
5h/day · 1 week
5h/week · 3 weeks
Not sure it’s for you? Come to the first session. If it’s not the right fit, we’ll help you find something that is.
AI Control 2