Skip to content
Tracks @XLab

The spillway routing mechanism

The spillway routing mechanism

reward hacking(misspecified reward)deceptive alignmentemergent misalignmentuncontrolled fitness-seekingspillway motivationdrives behavior“maximum score no matter what” → indifferentguides behaviorremaining motivations(plausibly aligned)

1. Training: misspecified reward

RL is intended to help the model learn useful skills, but the reward signal is sometimes misspecified, and the resulting training pressure causes harmful generalization.

Embed this demo

Use the chrome-less embed view inside an iframe:

<iframe src="/demos/spillway-routing/embed" width="100%" height="360" style="border:0"></iframe>