I think the Anthropic Risk Report that just came out is the most recent source for this, although I remember hearing similar claims floating around over the last year. A tweet about this: https://x.com/imjustnewatai/status/2088354827340296274
I think alignment faking might also be difficult because it involves training, which you can’t really do for frontier models (although maybe there are prompted versions which work similarly)
I think the Anthropic Risk Report that just came out is the most recent source for this, although I remember hearing similar claims floating around over the last year. A tweet about this: https://x.com/imjustnewatai/status/2088354827340296274
I think alignment faking might also be difficult because it involves training, which you can’t really do for frontier models (although maybe there are prompted versions which work similarly)