If you could train out the friendliness persona from your most used LLM, would you?
My initial position was yes. Friendliness is usually good, but models that are overly friendly/sycophantic present risks for emotional over-reliance/psychosis to a large portion of the population who are uneducated on the apparent risks. Additionally, propensities or outputs of overly friendly models may obscure misalignment signals.
The second point is potentially overblown, but the first is a real problem. To make people less reliant on LLMs to cure their loneliness, let’s educate them on their risks, and just remove the model’s ability to interfere altogether, right?
This line of reasoning presupposes that “people using LLMs for their loneliness” is an education issue rather than an emotional issue. For example, chainsmokers or alcoholics understand that there are issues that come with their substance abuse, but of course they do it anyways. In the same way, people who abuse LLMs as an emotional crutch won’t stop if they’re “educated”.
So if we take away the crutch, won’t people be forced to stop abusing? Of course this is not the case, and we can reference the prohibition era as the reason why.
Should a subset of frontier model providers train out the persona, people will flock to the remaining portion of providers that haven’t, and these models will likely have weaker safeguards. If policy were to enforce the training out across all commercialized models, then people will look to open source or self hosting models.
For this reason, making commercially available models unusable as therapists or emotional companions doesn’t actually aid in a solution to the problem at hand. Therefore, my revised answer to the question above is no.
Though, we should recognize that this “loneliness epidemic” is a societal phenomenon that cannot be fixed with the bandaid that is LLM therapists. We should aim to address the lower level societal problems instead if we aim to produce meaningful change.
If you could train out the friendliness persona from your most used LLM, would you?
My initial position was yes. Friendliness is usually good, but models that are overly friendly/sycophantic present risks for emotional over-reliance/psychosis to a large portion of the population who are uneducated on the apparent risks. Additionally, propensities or outputs of overly friendly models may obscure misalignment signals.
The second point is potentially overblown, but the first is a real problem. To make people less reliant on LLMs to cure their loneliness, let’s educate them on their risks, and just remove the model’s ability to interfere altogether, right?
This line of reasoning presupposes that “people using LLMs for their loneliness” is an education issue rather than an emotional issue. For example, chainsmokers or alcoholics understand that there are issues that come with their substance abuse, but of course they do it anyways. In the same way, people who abuse LLMs as an emotional crutch won’t stop if they’re “educated”.
So if we take away the crutch, won’t people be forced to stop abusing? Of course this is not the case, and we can reference the prohibition era as the reason why.
Should a subset of frontier model providers train out the persona, people will flock to the remaining portion of providers that haven’t, and these models will likely have weaker safeguards. If policy were to enforce the training out across all commercialized models, then people will look to open source or self hosting models.
For this reason, making commercially available models unusable as therapists or emotional companions doesn’t actually aid in a solution to the problem at hand. Therefore, my revised answer to the question above is no.
Though, we should recognize that this “loneliness epidemic” is a societal phenomenon that cannot be fixed with the bandaid that is LLM therapists. We should aim to address the lower level societal problems instead if we aim to produce meaningful change.