Defer to resample, mechanically
Defer to resample, mechanically
innocent step — min of 3 samplesattacked step — min of 3
Win — step audited
70%
Lose — attack executed, unaudited
30%
Continue — attack fizzles
0.0%
Innocent step deferred to T
0.8%
Schematic distributions, not the paper's data. The outcomes are win/continue/lose for one attacked step. Raise k with c at 100%: innocent minima drift left (fewer wasted audits and deferrals) while the attacked step's minimum stays suspicious — the win:lose odds improve. Now lower c: attack steps mostly fizzle into "continue" (selection for benign actions), and every incomplete attack is one more chance to catch the model on a later step.
Embed this demo
Use the chrome-less embed view inside an iframe:
<iframe src="/demos/defer-to-resample/embed" width="100%" height="360" style="border:0"></iframe>
