Skip to content

Defer to resample, mechanically

innocent step — min of 3 samplesattacked step — min of 3
defer →audit (top 1% of innocent minima) →more suspicious →

Win — step audited

70%

Lose — attack executed, unaudited

30%

Continue — attack fizzles

0.0%

Innocent step deferred to T

0.8%

Schematic distributions, not the paper's data. The outcomes are win/continue/lose for one attacked step. Raise k with c at 100%: innocent minima drift left (fewer wasted audits and deferrals) while the attacked step's minimum stays suspicious — the win:lose odds improve. Now lower c: attack steps mostly fizzle into "continue" (selection for benign actions), and every incomplete attack is one more chance to catch the model on a later step.