>If personas are a viable path to near-term alignment (and I think they are), control could set up a more adversarial relationship with the AI and increase the probability of misalignment that way.
I have some opinions on this (vibe based too):
if the AI is not very capable, then it could produce honest mistakes like https://incidentdatabase.ai/cite/1152/. Some control measures, like proper permission management, could help prevent this type of mistakes from escalating into disasters.
Even if control does more harm than good to overall safety of advanced AI, that doesn’t mean we should give it up now. We can use it now when it’s still beneficial, then drop it when it’s no longer so.
Overall I think this argument makes sense, and has some truth to it, but it doesn’t mean we should give up on control now.
I mean, if the adversarial relationship concern is true and significant, that is a good reason to just give up on it now before we make things worse.
In general, it seems bad to have a situation where an intelligence is trying or would want to be trying to harm us, even if we have measures which can prevent it from doing this. This is just not a very robust situation as intelligence scales.
>If personas are a viable path to near-term alignment (and I think they are), control could set up a more adversarial relationship with the AI and increase the probability of misalignment that way.
I have some opinions on this (vibe based too):
if the AI is not very capable, then it could produce honest mistakes like https://incidentdatabase.ai/cite/1152/. Some control measures, like proper permission management, could help prevent this type of mistakes from escalating into disasters.
Even if control does more harm than good to overall safety of advanced AI, that doesn’t mean we should give it up now. We can use it now when it’s still beneficial, then drop it when it’s no longer so.
Overall I think this argument makes sense, and has some truth to it, but it doesn’t mean we should give up on control now.
I mean, if the adversarial relationship concern is true and significant, that is a good reason to just give up on it now before we make things worse.
In general, it seems bad to have a situation where an intelligence is trying or would want to be trying to harm us, even if we have measures which can prevent it from doing this. This is just not a very robust situation as intelligence scales.