Someone can very reliably hurt/incapacitate people / fuck them up really hard, while also earnestly claiming and believing (or at least thinking that they believe) that all of this is either instrumental to their noble goal or an acceptable collateral damage of things that are instrumental to their noble goal.
ETA: What I mean is that the concept of “intrinsically valuing” diverges/splinters/[is no longer a good concept] when talking about the behavior of agents that is inadequately described in terms of “having beliefs” (including about what they “value”, what is “good”, etc.), which I think is the case here. Or, like, the player is very oblivious to what the character is doing. (See also: Enemies vs Malefactors.)
I’ll try to phrase my thoughts in a different way: in rationalist spheres, its’s extremely low social status to ponder “maybe X is just puppy kicking evil without any complex hidden upside or justification” and as a result this hypothesis doesn’t gain probability mass even when it makes good predictions.
After the Waluigi effect was discovered in LLMs, I had to reckon with the reality that minds exist which pretty well know what the right thing to do is and take the opposite action because it is opposite.
Lots of rationalist believe that death is bad period, without any nuance or benefits hiding in fine print in a vast majority of cases.
I’m failing to get your point about the Waluigi effect and how it connects to what you are talking about.
ETA: And maybe it’s sleep deprivation but I don’t see how this comment reframes the one I replied to, rather than talking about something entirely different.
What I think Hastings is getting at is that if we observe someone doing things that seem evil, we generally assume that inside their own world view they think what they’re doing is actually good. This assumption is because we have a “strong prior” that nobody is just willingly evil on purpose.
But if that assumption is too stong (a “locked prior”) it makes us vulnerable if we ever do encounter someone who’s just evil on purpose and knows it.
Personally I think Soryu is probably a case where he believes what he’s doing is good, but I think Hastings is cautioning against making that assumption too strongly in all cases.
I don’t know[1] how much the (most likely true) assumption that most people think they’re doing good actually buys us. People can have/develop very different conceptions of what Good is and what the instrumental strategies are that reliably get to the Good, sometimes so much that all hope of intelligibility is lost, especially when the person has twisted their environment, but also their own mind, into such a shape that questioning the load-bearing assumptions of what Good is and how to determine a Good strategy ~always fall flat.
Someone can very reliably hurt/incapacitate people / fuck them up really hard, while also earnestly claiming and believing (or at least thinking that they believe) that all of this is either instrumental to their noble goal or an acceptable collateral damage of things that are instrumental to their noble goal.
ETA: What I mean is that the concept of “intrinsically valuing” diverges/splinters/[is no longer a good concept] when talking about the behavior of agents that is inadequately described in terms of “having beliefs” (including about what they “value”, what is “good”, etc.), which I think is the case here. Or, like, the player is very oblivious to what the character is doing. (See also: Enemies vs Malefactors.)
I’ll try to phrase my thoughts in a different way: in rationalist spheres, its’s extremely low social status to ponder “maybe X is just puppy kicking evil without any complex hidden upside or justification” and as a result this hypothesis doesn’t gain probability mass even when it makes good predictions.
After the Waluigi effect was discovered in LLMs, I had to reckon with the reality that minds exist which pretty well know what the right thing to do is and take the opposite action because it is opposite.
Lots of rationalist believe that death is bad period, without any nuance or benefits hiding in fine print in a vast majority of cases.
I’m failing to get your point about the Waluigi effect and how it connects to what you are talking about.
ETA: And maybe it’s sleep deprivation but I don’t see how this comment reframes the one I replied to, rather than talking about something entirely different.
What I think Hastings is getting at is that if we observe someone doing things that seem evil, we generally assume that inside their own world view they think what they’re doing is actually good. This assumption is because we have a “strong prior” that nobody is just willingly evil on purpose.
But if that assumption is too stong (a “locked prior”) it makes us vulnerable if we ever do encounter someone who’s just evil on purpose and knows it.
Personally I think Soryu is probably a case where he believes what he’s doing is good, but I think Hastings is cautioning against making that assumption too strongly in all cases.
Big thanks for the clarification.
I don’t know[1] how much the (most likely true) assumption that most people think they’re doing good actually buys us. People can have/develop very different conceptions of what Good is and what the instrumental strategies are that reliably get to the Good, sometimes so much that all hope of intelligibility is lost, especially when the person has twisted their environment, but also their own mind, into such a shape that questioning the load-bearing assumptions of what Good is and how to determine a Good strategy ~always fall flat.
When I say I don’t know, I mean that I don’t know, rather than that I very strongly doubt it.