Courses/AI Control

AI Control

How can we safely use powerful AI systems that may actively subvert oversight? Study threat models, control evaluations, monitoring, and practical protocols.

Enroll for free
Next cohortDates TBA

Leave your email and we’ll let you know.

Pace
Mon–Fri or weekly
Per unit
~5h
Format
Online, small groups
Included
1-1 advising call
Certificate
Tiered, verifiable

About this course

AI control asks a pragmatic question: if a powerful system may be misaligned and may deliberately undermine oversight, can we still use it without giving it a path to catastrophe? You’ll learn how control protocols combine trusted models, untrusted models, monitors, resampling, audits, and carefully designed environments.

This is a hands-on technical course. You’ll build attack trees, reason about safety–usefulness tradeoffs, evaluate high-stakes deployment protocols, and investigate diffuse failures such as research sabotage, sandbagging, and capability-elicitation failures. Exercises and small-group discussion turn the field’s primary literature into decisions you can defend.

The course assumes introductory AI safety knowledge and comfort reading ML or AI safety papers and technical blog posts. You do not need prior AI control research experience; the course builds the threat models and vocabulary before asking you to design or critique a protocol.

Who is it for:

ML engineers

You want to apply practical engineering skills to monitoring, evaluation, agent environments, or safer deployment protocols.Build systems for monitoring and safer deployment.

Technical researchers

You want a structured route through the core control literature and a sharper view of the open empirical questions.Turn the control literature into research questions.

Evaluators and red-teamers

You study model failures and want to connect threat models, capability elicitation, sandbagging, and sabotage to concrete evaluations.Connect threat models to concrete evaluations.

AI safety strategists

You want to judge when control can reduce risk, how much useful work a protocol preserves, and where control fits beside alignment and governance.Judge when control helps and where it fits.

Curriculum

5 UNITS · ~5H EACH

Adapted from XLab’s seven-module AI Safety Tracks curriculum. Lens groups the primary readings and exercises into five cohort units, with facilitated discussion and assessed outputs.

01

Threat models for intentionally subversive AI

Make the case for control without assuming it works. Build attack trees, distinguish concentrated from diffuse failures, and prioritize the paths by which a powerful model could undermine its operators.

02

Safety–usefulness tradeoffs and a taxonomy of control

Model the frontier between useful work and acceptable risk. Compare detection, prevention, and response mitigations; then connect them to control evaluations and the environments in which protocols are tested.

03

High-stakes control

Study win, continue, and lose outcomes; execute, replace, and audit protocols; trusted and untrusted monitoring; resampling; agent-system architecture; and the practical challenges of collusion and legibility.

04

Sabotage, sandbagging, and capability elicitation

Move beyond single catastrophic actions to failures spread across many ordinary tasks. Examine research sabotage, strategic underperformance, exploration hacking, and whether evaluations elicit the capabilities a model may hide.

05

Beyond schemers and toward a research agenda

Consider reward-seeking systems, spillway motivations, cheaply satisfied preferences, and possible deals with early misaligned systems. Finish by selecting a tractable control question and designing a research or evaluation plan.

Course Outcomes

A defensible control threat model

An attack tree that names the system, oversight boundary, failure paths, and assumptions your protocol must survive.

A safety–usefulness model

A way to compare protocols by the risk they remove and the useful work they preserve, rather than by safety claims alone.

A control protocol you can critique

A concrete combination of monitoring, auditing, resampling, access limits, and response rules, with its residual risks made explicit.

Stronger evaluation instincts

The ability to spot sandbagging, elicitation failures, collusion risks, and benchmarks that test the wrong threat model.

Fluency in the core literature

Enough context to follow current AI control research and explain how high-stakes and diffuse-failure work differ.

A concrete next research step

An experiment, evaluation proposal, or research question narrow enough to investigate after the course.

Upcoming cohorts

INTENSIVE

5h/day · 1 week

Dates to be announcedWe’ll publish the next intensive cohort here.

PART-TIME

5h/week · 5 weeks

Dates to be announcedWe’ll publish the next part-time cohort here.

Common questions

Not sure it’s for you? Come to the first session. If it’s not the right fit, we’ll help you find something that is.

Get a 1-1 advising call with someone working on AI safety full-time.

Enroll for free

AI Control

Enroll free