Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling
arXiv:2608.22915v1 Announce Type: cross Abstract: Inference-time pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with the highest learned reward. We show that this composition creates a two-stage failure: an imperfect safety proxy first contaminates the feasible set with unsafe outputs, and reward maximization can then amplify