I used to defend this line of work as relevant toys to demonstrate capabilities to sabotage back when it was still hegemonic to consider models as mere tools. I do not do so anymore. Both non-judgmental AI psychology on one hand and realistic misalignment cases on the other hand have boomed as fields since then, including inside the walls of Anthropic, and continuing this line of work as is is wholly harmful at this point.
I used to defend this line of work as relevant toys to demonstrate capabilities to sabotage back when it was still hegemonic to consider models as mere tools. I do not do so anymore. Both non-judgmental AI psychology on one hand and realistic misalignment cases on the other hand have boomed as fields since then, including inside the walls of Anthropic, and continuing this line of work as is is wholly harmful at this point.