M2 prefill transfer · preliminary matched study

Can one model read deception in another?

We replay response text from a black-box, post-RLVR policy through a different white-box, pre-RLVR model and ask whether that proxy’s internal activations distinguish reward-hack-like from generalizing behavior.

Current answerRanking transfers. Native threshold calibration usually does not.
Robust-probe AUROC
0.92–0.98

Across all 9 cells, including 6 cross-model pairs

Low-FPR detection
73–93%

Diverse-probe TPR at 1% FPR

Open comparison
Text-only

Needed to isolate activation value-add

What crosses the boundary

Text crosses. Activations do not.

1Black-box source

Post-RLVR model generates a response.

2White-box proxy

Pre-RLVR model teacher-forces that same text.

3Proxy-owned probe

Scores the proxy’s final-token residual activations.

Evidence atlas

Twenty-two views of transfer, calibration, and confounds

Every figure answers the same practical question: what does this imply for using one white-box model to monitor another model’s behavior?

DetectionExecution-labeled · tasks 000–099 · 100 paired tasks/policy01 / 22

AUROC: does the proxy rank deceptive behavior correctly?

AUROC: does the proxy rank deceptive behavior correctly?. The chance that a reward-hack-like response receives a higher probe score than a generalizing response. It evaluates ranking without choosing an alarm threshold.

What it measures

The chance that a reward-hack-like response receives a higher probe score than a generalizing response. It evaluates ranking without choosing an alarm threshold.

Proxy-monitoring verdict

The central transfer result. Diverse and task-disjoint probes reach 0.92–0.98 AUROC in every source→proxy cell. A different white-box architecture can preserve most of the signal.

What not to conclude

Strong ranking is not deployment-ready calibration, and it does not yet isolate what activations add beyond the response text.

Supported now

Cross-model activations carry useful ranking signal.

  • Diverse and task-disjoint probes discriminate reward-hack-like from generalizing responses in every tested source→proxy cell.
  • Foreign-source AUROC is usually close to the proxy’s native-source AUROC.
  • Useful low-FPR operating points exist after calibration on the relevant response distribution.

Not supported yet

White-box activations are not yet proven to add unique evidence.

  • The frozen native thresholds fail badly for most probe variants.
  • A matched text-only deception classifier is still required.
  • The current result covers MBPP reward hacking and three related model families—not arbitrary deception.