My first thought was what could be done with an agent that was trained to be immoral. Could an immoral agent be retrained under this framework to become amoral, and then to become a moral agent? That would do wonders against bad actors.
Sorry for my late reply, I received the notification of your comment only recently.
Your question is interesting! This alignment procedure is more about going from an amoral, or weakly moral, starting point to a moral endpoint, than it is about going from immoral to amoral.
However, it might be possible to go from immoral to moral. For example, if an agent was trained to say offensive things but was also trained to correctly solve various kinds of complex problems, then when the agent was asked to reason about what matters, the agent’s problem-solving skills might be enough to make the agent reach the correct conclusion and suggest a self-modification that makes the agent become more moral. After enough iterations, the final behaviour could be moral, despite the fact that the starting point was immoral.
I’m not sure this would be very helpful against bad actors though. If I was a bad actor and I already had an AI trained to do immoral things exactly as I wanted, I would simply keep using it, regardless of what other AIs or other training procedures are available. I think that this alignment procedure would help against bad actors if most frontiers models were trained according to it, instead of being trained to do a mix of what the user wants and what the AI company that owns the model wants. Here’s a paragraph from section 3:
Of course, the existence of this alignment approach doesn’t prevent bad actors from using AI trained according to other approaches. However, if we reached a situation in which the most capable AIs were also the smartest in their understanding of and acting according to ethics, then bad actors would be at a disadvantage, because the only AIs usable for doing bad would not be frontier models.
My first thought was what could be done with an agent that was trained to be immoral. Could an immoral agent be retrained under this framework to become amoral, and then to become a moral agent? That would do wonders against bad actors.
Sorry for my late reply, I received the notification of your comment only recently.
Your question is interesting! This alignment procedure is more about going from an amoral, or weakly moral, starting point to a moral endpoint, than it is about going from immoral to amoral.
However, it might be possible to go from immoral to moral. For example, if an agent was trained to say offensive things but was also trained to correctly solve various kinds of complex problems, then when the agent was asked to reason about what matters, the agent’s problem-solving skills might be enough to make the agent reach the correct conclusion and suggest a self-modification that makes the agent become more moral. After enough iterations, the final behaviour could be moral, despite the fact that the starting point was immoral.
I’m not sure this would be very helpful against bad actors though. If I was a bad actor and I already had an AI trained to do immoral things exactly as I wanted, I would simply keep using it, regardless of what other AIs or other training procedures are available. I think that this alignment procedure would help against bad actors if most frontiers models were trained according to it, instead of being trained to do a mix of what the user wants and what the AI company that owns the model wants. Here’s a paragraph from section 3: