Skip to content

Control is additive

Step through the additive property of control & note how alignment interventions act on the model in an inner box, control's applied monitoring wraps around them, and toggling the shell leaves the inner box untouched.

Control — applied monitoringtrusted monitoringauditingdefer-to-trustedresamplingAlignment interventionsact on the model itselftraining interventionsclassifiers / probes

1. Alignment interventions

Training interventions (RLHF, adversarial training) and classifiers or probes act on the model itself, trying to make it trustworthy.