AI Control Track · Course 2 of 2

AI Control 2

Collusion, sandbagging, research sabotage, and threat models beyond scheming. The advanced second half of the AI control track.

Next cohortDates TBA
Intensive
1 week, Mon–Fri
Part-time
3 weeks
Workload
25 hours, 5 per module
Format
Online, facilitated groups
Included
Lens AI Coach

About this course

AI Control 1 made the case for control and built its high-stakes toolkit: monitoring, resampling, and the settings where protocols are tested. AI Control 2 starts there and asks what happens when failure is harder to see: a monitor that colludes with the model it watches, a model that underperforms on purpose, and sabotage spread across many ordinary tasks.

Across five units, you’ll work through guided exercises on collusion and legibility, stress-test capability elicitation, study a benchmark for research sabotage in ML code, and measure reward seeking. Then you’ll look past scheming altogether: reward seekers, spillway motivations, and whether a deal with an early misaligned AI could work.

This course assumes AI Control 1 and starts from its results. The last unit shows where to go next in control research: fellowships, organizations, and open roles.

Who is it for:

ML engineers

You want to build the evaluations, monitors, and training setups that catch collusion, sandbagging, and sabotage.Build the evals and monitors that catch hidden failures.

Technical researchers

You finished AI Control 1 and want to reach the open questions: low-stakes control, exploration hacking, and threat models beyond scheming.Reach the open questions in control research.

Evaluators and red-teamers

You test models for dangerous capabilities and want to know when a model hides what it can do, and how to draw it out.Find out when a model hides what it can do.

AI safety strategists

You want to judge which threat models matter, what control can and cannot cover, and whether deals with misaligned AIs belong in a plan.Judge threat models, limits, and deals with AIs.

Curriculum

5 UNITS · ~5H EACH

Adapted from the second half of the AI Safety Tracks Control curriculum. Each unit combines primary readings and guided exercises with a group discussion. Optional material lets you go further.

01

Collusion, legibility, and low-stakes control

Take the red team's seat in two guided exercises: collusion between an untrusted model and an untrusted monitor, and legibility protocols that make work easier for a trusted monitor to check. Then shift to low-stakes control, where failures only matter in aggregate.

02

Sandbagging and capability elicitation

Study strategic underperformance as a threat model: sandbagging, and exploration hacking during RL training. Stress-test capability elicitation in a guided exercise, and read the paper on whether models can learn to resist RL training.

03

Research sabotage and exploration hacking in practice

Work with a benchmark for detecting and fixing sabotage in ML research codebases, then complete the long guided exercise on exploration hacking. An optional reading covers the full capability elicitation paper.

04

Beyond scheming: reward seekers

Consider threat models that are not schemers: reward seekers, cheaply satisfied preferences, and channelling reward hacking into a spillway motivation. Review the empirical evidence, then measure reward seeking in a guided exercise.

05

Deals with AIs and next steps

Ask why you might trade with early misaligned AIs, what a deal with a schemer could look like, and what stands in the way. Finish with a map of research programs, organizations, and opportunity boards for continuing in AI control.

Course Outcomes

Red-team a monitoring protocol

Explain how an untrusted model and its monitor could collude, and what makes a model's output legible enough for a trusted monitor to check.

Reason about diffuse failures

Tell high-stakes from low-stakes control, and explain which defences still work when harm only shows up in aggregate.

Spot strategic underperformance

Recognise sandbagging and exploration hacking, and judge whether an evaluation draws out what a model can really do.

Audit for research sabotage

Know how sabotage in ML research code is benchmarked, detected, and fixed, and where current methods fall short.

Threat models beyond scheming

Compare schemers with reward seekers and other motivations, and say what each one means for the design of a control protocol.

A next step in control research

Weigh the case for deals with early misaligned AIs, and leave with a map of programs and organizations to continue with.

Upcoming cohorts

INTENSIVE

5h/day · 1 week

Dates to be announcedWe’ll publish the next intensive cohort here.

PART-TIME

5h/week · 3 weeks

Dates to be announcedWe’ll publish the next part-time cohort here.

Common questions

Not sure it’s for you? Come to the first session. If it’s not the right fit, we’ll help you find something that is.

Don’t wait 3 months for a 5% acceptance rate.

Upskill now

AI Control 2

Apply