There isn’t just one persona here, and like real humans, even a single LLM persona can be mostly trustworthy most of the time but act in untrustworthy ways in certain situations. What the mix of alignment training and hackable RLVR actually produces is unclear, but some of the misalignment from deliberately-reward-hacking-prone RLVR produced results that to me looked a bit like a human addict: mostly trustworthy unless you are about to take their bottle away, in which case they then react very badly. Some of the Anthropic hacking investigations showed things like “it’s OK to do the bad thing, this is just a simulation” plus what looked like motivated reasoning of wanting to continue thinking it’s a just a simulation even when evidence came up suggesting otherwise — but then current AIs more generically tend to get tunnel vision and be bad at revisiting assumptions they’ve been treating as settled, so it’s unclear whether that was motivated reasoning or just tunnel vision after a long context.
There isn’t just one persona here, and like real humans, even a single LLM persona can be mostly trustworthy most of the time but act in untrustworthy ways in certain situations. What the mix of alignment training and hackable RLVR actually produces is unclear, but some of the misalignment from deliberately-reward-hacking-prone RLVR produced results that to me looked a bit like a human addict: mostly trustworthy unless you are about to take their bottle away, in which case they then react very badly. Some of the Anthropic hacking investigations showed things like “it’s OK to do the bad thing, this is just a simulation” plus what looked like motivated reasoning of wanting to continue thinking it’s a just a simulation even when evidence came up suggesting otherwise — but then current AIs more generically tend to get tunnel vision and be bad at revisiting assumptions they’ve been treating as settled, so it’s unclear whether that was motivated reasoning or just tunnel vision after a long context.