Lifelong recursive self-improver, on his way to exploding really intelligently :D
More seriously: my posts are mostly about AI alignment, with an eye towards moral progress. I have a bachelor’s degree in mathematics, I did research at CEEALAR for four years, and now I do research independently.
Imagine it’s the year 1500, but AI technology is available. You want to make an AI that can tell you that witch hunts are a terrible idea and can convincingly explain why, despite the fact that many people around you seem to think the exact opposite. How do you do it?
The problem above, and the fact that I’d like to avoid producing AI that can be used for bad purposes, is what motivates my research. I think I’ve made progress towards a solution, in both theoretical and practical terms. Regarding why this matters, these two short posts are a good starting point. If you are looking for something more technical, consider setting some time aside to read these two.
Feel free to reach out!
You can support my research through Patreon here.
Sorry for my late reply, I received the notification of your comment only recently.
Your question is interesting! This alignment procedure is more about going from an amoral, or weakly moral, starting point to a moral endpoint, than it is about going from immoral to amoral.
However, it might be possible to go from immoral to moral. For example, if an agent was trained to say offensive things but was also trained to correctly solve various kinds of complex problems, then when the agent was asked to reason about what matters, the agent’s problem-solving skills might be enough to make the agent reach the correct conclusion and suggest a self-modification that makes the agent become more moral. After enough iterations, the final behaviour could be moral, despite the fact that the starting point was immoral.
I’m not sure this would be very helpful against bad actors though. If I was a bad actor and I already had an AI trained to do immoral things exactly as I wanted, I would simply keep using it, regardless of what other AIs or other training procedures are available. I think that this alignment procedure would help against bad actors if most frontiers models were trained according to it, instead of being trained to do a mix of what the user wants and what the AI company that owns the model wants. Here’s a paragraph from section 3: