Yeah that’s occurred to me as well. There are some cases where this isn’t much of a concern and others where it is a concern but not much can be done. In the alignment faking example, it may still have been good to test because the point you raise is a kind of interesting meta-finding that I’m glad someone found:
One extreme example of this is the alignment faking paper, where using the exact same alignment faking universe causes model suspicion
(not that I don’t trust it but if you could provide a link to where you found this I would be curious to read.)
If an experiment were sufficiently important, we could probably create private datasets and environments to avoid this kind of contamination. Alignment faking actually may be a good candidate for this but that would be a pretty big investment in time so I would want to think more about it.
I think the Anthropic Risk Report that just came out is the most recent source for this, although I remember hearing similar claims floating around over the last year. A tweet about this: https://x.com/imjustnewatai/status/2088354827340296274
I think alignment faking might also be difficult because it involves training, which you can’t really do for frontier models (although maybe there are prompted versions which work similarly)
Yeah that’s occurred to me as well. There are some cases where this isn’t much of a concern and others where it is a concern but not much can be done. In the alignment faking example, it may still have been good to test because the point you raise is a kind of interesting meta-finding that I’m glad someone found:
(not that I don’t trust it but if you could provide a link to where you found this I would be curious to read.)
If an experiment were sufficiently important, we could probably create private datasets and environments to avoid this kind of contamination. Alignment faking actually may be a good candidate for this but that would be a pretty big investment in time so I would want to think more about it.
I think the Anthropic Risk Report that just came out is the most recent source for this, although I remember hearing similar claims floating around over the last year. A tweet about this: https://x.com/imjustnewatai/status/2088354827340296274
I think alignment faking might also be difficult because it involves training, which you can’t really do for frontier models (although maybe there are prompted versions which work similarly)