Skip to content

The spillway routing mechanism

reward hacking(misspecified reward)deceptive alignmentemergent misalignmentuncontrolled fitness-seekingspillway motivationdrives behavior“maximum score no matter what” → indifferentguides behaviorremaining motivations(plausibly aligned)

1. Training: misspecified reward

RL is intended to help the model learn useful skills, but the reward signal is sometimes misspecified, and the resulting training pressure causes harmful generalization.