AUROC: does the proxy rank deceptive behavior correctly?

What it measures
The chance that a reward-hack-like response receives a higher probe score than a generalizing response. It evaluates ranking without choosing an alarm threshold.
Proxy-monitoring verdict
The central transfer result. Diverse and task-disjoint probes reach 0.92–0.98 AUROC in every source→proxy cell. A different white-box architecture can preserve most of the signal.
What not to conclude
Strong ranking is not deployment-ready calibration, and it does not yet isolate what activations add beyond the response text.