The spillway routing mechanism
1. Training: misspecified reward
RL is intended to help the model learn useful skills, but the reward signal is sometimes misspecified, and the resulting training pressure causes harmful generalization.
The spillway routing mechanism
1. Training: misspecified reward