Before the session
- Reread the strategic underperformance lesson and the setup of the password-locked models paper.
- Try the locking flowcharts exercise yourself; you will run it as a group.
- If you run a classroom, check the dashboard for who reached the capability-elicitation walkthrough; seat at least one finisher in each debate team.
Agenda
Times are a running clock for a 60-minute session
Warm-up: many small actions (0:00–0:10)
Write the question on the whiteboard: "What happens when no single action is catastrophic?" Collect answers unfiltered, then check the list against the opening of the non-concentrated failures lesson — star what the room got right, add what it missed.
Prompt:
- "Auditing every suspicious action worked in the high-stakes picture. Why might it not work here?"
Naming the failure modes (0:10–0:20)
Rapid definitions round: sandbagging, exploration hacking, password-locked models. A volunteer defines each in one or two sentences; the room refines until nobody objects, and you write the final versions up. Anchor on the strategic underperformance lesson and the password-locked models paper.
Prompts:
- "How do you tell sandbagging from honest low capability, from the outside?"
- "Why do the papers bother building password-locked models rather than waiting to find sandbagging in the wild?"
Flowcharts, together (0:20–0:35)
Run the locking flowcharts exercise on the shared screen, stage by stage: hand-count vote at each branch, minority explains its branch before you submit and grade. Revisit the branch the room got most wrong at the end.
Prompts:
- "What fact would you need to observe to be confident in this branch?"
The elicitation debate (0:35–0:55)
Two teams, five minutes preparation. One defends the claim that fine-tuning-based elicitation, as studied in the password-locked models paper, reliably defeats sandbagging; the other builds the strongest case that it fails, drawing on the exploration hacking paper and the elicitation walkthrough. Four minutes per side uninterrupted, then open floor with the facilitator only enforcing turn-taking. Close by asking which experiment from the auditing benchmark paper or the papers above would move the debate most.
Prompts:
- "What does elicitation need that a sandbagging model can withhold?"
- "If elicitation reliably works, what does a lab no longer need to worry about? If it fails, what breaks?"
Wrap-up (0:55–1:00)
Each participant states one way the low-stakes regime changed what they think monitoring is for. Collect open questions.
After the session
- On the classroom dashboard, check progress into the module's four papers; assign each debate team one paper to finish before next session.
- Read written submissions for the module; recurring confusions about sandbagging versus low capability are the usual warm-up for the next meeting.

