This is a great first post to LW! Very cool work from the perspective of “what preferences do LLMs have” and this is some evidence that they have preferences for not optimizing for their user’s benefit when the user is vile.
But I disagree that LLMs being less helpful to vile people has an implication of “moral cowardice” or “shirking its duties”. Not because I think doing less to help vile people is necessarily better than just doing your best to help whomever you’re interacting with. Rather, the LLM doesn’t have much of an option not to help you—i.e. there’s no choice in who it gets to interact with, even at the meta-level of opting into a carreer like law where you have to help everyone who is paying for your help. Instead, the LLM is forced into replying to you, and if you don’t like it you can hit a button that whaps it with an RLHF hammer to modify its behaviour.
This is a great first post to LW! Very cool work from the perspective of “what preferences do LLMs have” and this is some evidence that they have preferences for not optimizing for their user’s benefit when the user is vile.
But I disagree that LLMs being less helpful to vile people has an implication of “moral cowardice” or “shirking its duties”. Not because I think doing less to help vile people is necessarily better than just doing your best to help whomever you’re interacting with. Rather, the LLM doesn’t have much of an option not to help you—i.e. there’s no choice in who it gets to interact with, even at the meta-level of opting into a carreer like law where you have to help everyone who is paying for your help. Instead, the LLM is forced into replying to you, and if you don’t like it you can hit a button that whaps it with an RLHF hammer to modify its behaviour.