Thanks to Dennis Akar, Rauno Arike, Shubhorup Biswas, Claude Fable, Max Heitmann, Vladimir Ivanov, and Jordan Taylor (alphabetical order) for helpful comments on this draft and on the research so far. Thanks to Rohan Subramani and Rhys Ward for high-level comments and discussion. Based on project proposals from Max Heitmann, Jordan Taylor, and Joshua Clymer. This work was done while at Aether Research.
Code available here, metrics & run info available here.
TL;DR: Held-out evals / monitors / probes would be really nice to have, but the “held-out-ness” is easier claimed than guaranteed. We measure a generalized form of “feedback spillover” and show that training against an LLM monitor can sometimes degrade a deception probe, and vice versa.
Executive Summary
* It seems crucial to have measures of alignment that still work, even though we train on other measures of alignment. Whether we get this by default is an open question.
* We run preliminary experiments on a suite of probes and LLM monitors, and report the following:
* Training against one proxy can produce reward hacking policies that are less suspicious.
* Proxies also become worse at discriminating hacks from non-hacks, even when not trained against.
* We can observe the correlated degradation of different proxies, and note that on the occasions when a strong monitor is evaded, other monitors are evaded even more strongly.
* Our preliminary recommendation: if you are using a held-out proxy to evaluate your model, you should have empirical or theoretical reasons to assume it won’t degrade during training.
* A science of held-out monitors seems possible; we currently don’t have one. This makes safety cases that rely on held-out alignment proxies less reassuring than hoped.
Introduction
It seems like the default “alignment plan” is going to be haphazard, relying on defense-in-depth to make up for a lack of fundamental breakthroughs. In these worlds, we’ll often have to train on proxies for the