Skip to content
Tracks @XLab

Exercises

Lorem ipsum — standalone exercises to work through in place; the same cards embed inside lessons and papers.

Judgment as to the Usefulness of Control
Allocate a team of 10 researchers across five safety agendas, once per scenario, across four scenarios.
OpenAllocation exercise
Responding to Arguments Against Control
Respond to criticisms of AI control, then see one response defenders give. Ends by building your own argument.
OpenArgument exercise
Run the evaluation — 5 scenarios
Five deployment setups, the same question each time: assuming the untrusted model is scheming, does the setup hold the control property?
OpenScenario exercise
Why catching counts, and where evaluation runs out
Two reasoning prompts on the game theory of control: why a catch changes the game, and where capability evaluations stop being informative.
OpenReasoning prompts
Writing prompt
Pick one area of control work and construct an explicit causal chain from work starting today to reduced existential risk.
OpenWriting prompt
Writing prompt
Find the link in your theory-of-change chain you trust least, and name the observations that would raise or lower your confidence.
OpenWriting prompt
Build the flow chart
Reconstruct the password-locked model construction as three drag-and-drop flow charts: strong policy, weak policy, locked model.
OpenBuild the flow chart
Memo
Draft the terms of a deal a lab could offer a model it suspects of sabotage: task, compensation, delivery, and verification.
OpenMemo