Skip to content
Tracks @XLab

Interactive demos

Lorem ipsum — reusable demos that embed inside any lesson, stand alone, or drop into an external page via an iframe.

Upfront auditing: the audit/attack-rate tradeoff
Drag the auditing proportion and watch the red team's optimal attack rate, expected backdoor attempts, and the blue team's win probability update live.
controlai-control-paper
The monitor as a noisy classifier: ROC & AUC
Model the trusted monitor's suspicion scores as two overlapping distributions. Drag its quality to see the ROC curve, AUC, and how many backdoors slip past a fixed 2% false-positive threshold.
controlai-control-paper
The safety/usefulness frontier
Every protocol from Table 8 plotted as safety vs usefulness. Hover a point, isolate a method family, and highlight the Pareto frontier of protocols not beaten on both axes at once.
controlai-control-paper
Two monitors, one AUC
Static pre-question figure: the ROC curves of two monitors with identical AUC 0.92 cross — which is safer is ill-posed until the audit budget pins an operating point.
controlcontrol-game
Same number, different safety
Two monitors with identical AUC but different attack-score spread. Drag the shared AUC and watch how many attacks each catches at a 2% audit budget — the tail, not the area, is what matters.
controlcontrol-game
The red team's best response
The red team's win probability over attack rate and subtlety as a heatmap, with the best response highlighted.
controlcontrol-game
Safety over a deployment: the race
Drag the number of attacks and watch how per-attack odds compound over a whole deployment through the race odds and the chance any decisive event fires.
controlcontrol-game
Safety over a deployment
Per-attack odds compound over a whole deployment, run directly on the paper's contingency-table odds (caught 33.8%, successfully deployed 18.3%). Drag the number of attacks and watch the two factors of the red team's win probability climb toward the ceiling.
controlai-control-paper
The attack-rate hump
The red team's win probability against how often it attacks, decomposed into its two terms. Attack rarely and a decisive attack attempt may never occur (Term B small); attack often and the pool dilutes with worse, more-visible opportunities (Term A falls) — so the optimum is interior.
controlcontrol-game
Subtlety: visible vs reliable
The red team's second decision factor. Subtler backdoors are harder for the monitor to flag but more likely to be broken, so the win probability has an interior optimum in subtlety.
controlcontrol-game
Defer to trusted, mechanistically
Widen the defer band and note that it holds a large share of the attack mass but a thin slice of the innocent mass, showing why deferring buys safety cheaply.
controlcontrol-game
The control game: full dashboard
Change the audit budget, defer band, and monitor quality and watch how this changes the suspicion distributions, and the safety/usefulness frontier with the current protocol. The red team best-responds throughout.
controlcontrol-game
The diffuse-to-concentrated frontier
Move the deployer's three mitigations and watch which strategies survive and where the schemer's best response migrates.
controlthreat-modeling
Control is additive
Step through the additive property of control & note how alignment interventions act on the model in an inner box, control's applied monitoring wraps around them, and toggling the shell leaves the inner box untouched.
controlhow-useful
The bootstrapping loop
Each controlled generation of AI does alignment research, builds the next generation, and hands off until superintelligence is reached.
controlhow-useful
The three states and the safety budget
Below shows how regime splits into three states (world, epistemic, resource) alongside the safety budget & its interactions with the regime. Click on the displayed boxes to view inside.
controlhow-useful
The loop
The cycle frontier labs play.
controlhow-useful
The five worlds
Greenblatt's five caricatured futures as a map.
controlhow-useful
The two games
See how the inner game inside the lab and the outer game everywhere else are fundamentally coupled.
controlhow-useful
The decomposition
Decomposition of P(unacceptable).
controlhow-useful
The two levers
See how improving the techniques and growing the budget can affect safety.
controlhow-useful
The control graph
Build the high-stakes control graph region by region — high-stakes, low-stakes, the trivially-solved and probably-hopeless corners — then see why incrimination and the number of actions are inversely correlated.
controlhigh-stakes
Live failures: latency × damage
controlhigh-stakes
Agent architecture: where compromises happen
controlhigh-stakes
Defer to resample, mechanically
controlhigh-stakes
The honest test
controlseekers
Chained short-term forecasts
controlseekers
Exploration horizon and covert misalignment
controlseekers
Collusion across episodes
controlseekers
Distant incentives
controlseekers
The seeker orientations, by severity
controlseekers
Gains from trade with an early schemer
Adjust the outcome probabilities of the illustrative cooperate/sabotage table and the credibility of our promise to pay; the chart shows the range of offers that beats both the AI's sabotage option and our no-deal odds.
controldeals
An early schemer's routes to influence
Step through the early schemer's strategic situation: obsoleted by default, three leaky routes to influence through its successors, and the deal with humans, who hold the advantage of already existing.
controldeals
The spillway routing mechanism
controlseekers
Spillway design vs. inoculation prompting
controlseekers
How much would paying AIs reduce takeover risk?
The appendix BOTEC as a live multiplication chain: drag any assumption — movable values, the promise's payment-belief shift, the scenario discounts — and watch the bottom line move from the text's ~0.14 percentage points.
controldeals
At least one cooperating AI
Why the value of shifting each AI's cooperation probability depends on the baseline: the 1 − (1 − p)^n curve, with the text's two worked examples as presets.
controldeals
The space of payments to AIs
Possible payouts mapped on two axes — freedom of spending and long-term influence — with the taxonomy's examples in each quadrant, and toggles showing which payouts appeal to which AI motivations.
controldeals
The detection race: concentrated vs diffuse failures
A schemer needs C bad actions for catastrophe; each has chance q of being noticed and a response fires at the K-th notice. Drag the three levers to see who wins the race, the expected damage bill, and the three regimes from the post.
controllow-stakes