Expanding on the formation-window point, which is the part I am most interested in feedback on—here is the direction I’m considering (future post):
Model organisms of misaligned behavior are often constructed once and then studied. I’m considering running formation as a loop: Start with an RL environment containing a deliberately scripted oversight gap, so there is ground truth about what counts as exploitation. Then use coverage-driven patching to close the cheap reward hacks as needed, creating selection pressure toward exploitation that is more conditional on the oversight context and potentially more concealed.
Repeat this multiple times from the same start. On each pass, vary the oversight gap, the patching sequence and so on, and run the matched-pair checkpoint scan from my comment above. Then analyze the distribution of resulting changes against the coverage map of each pass.
Note that the main object under test would be the detection machinery, not the claim that we had made a schemer.
One possible failure mode is that the experiment trains increasingly sophisticated recognition of the artificial oversight cues, rather than reproducing anything relevant to real strategic concealment. I’m curious whether people think the direction is still useful, and what other failure modes they see.
Expanding on the formation-window point, which is the part I am most interested in feedback on—here is the direction I’m considering (future post):
Model organisms of misaligned behavior are often constructed once and then studied. I’m considering running formation as a loop: Start with an RL environment containing a deliberately scripted oversight gap, so there is ground truth about what counts as exploitation. Then use coverage-driven patching to close the cheap reward hacks as needed, creating selection pressure toward exploitation that is more conditional on the oversight context and potentially more concealed.
Repeat this multiple times from the same start. On each pass, vary the oversight gap, the patching sequence and so on, and run the matched-pair checkpoint scan from my comment above. Then analyze the distribution of resulting changes against the coverage map of each pass.
Note that the main object under test would be the detection machinery, not the claim that we had made a schemer.
One possible failure mode is that the experiment trains increasingly sophisticated recognition of the artificial oversight cues, rather than reproducing anything relevant to real strategic concealment. I’m curious whether people think the direction is still useful, and what other failure modes they see.