Skip to content
Tracks @XLab

Control is additive

Step through the additive property of control & note how alignment interventions act on the model in an inner box, control's applied monitoring wraps around them, and toggling the shell leaves the inner box untouched.

Control is additive

Step through the additive property of control & note how alignment interventions act on the model in an inner box, control's applied monitoring wraps around them, and toggling the shell leaves the inner box untouched.

Control — applied monitoringtrusted monitoringauditingdefer-to-trustedresamplingAlignment interventionsact on the model itselftraining interventionsclassifiers / probes

1. Alignment interventions

Training interventions (RLHF, adversarial training) and classifiers or probes act on the model itself, trying to make it trustworthy.

Embed this demo

Use the chrome-less embed view inside an iframe:

<iframe src="/demos/additive-control/embed" width="100%" height="360" style="border:0"></iframe>