Yes. But I think just a little fine tuning on selected notes or CoT makes Claude mostly analogous to you.
This is an interesting positive aspect of the way that LLM AGI will have memory, and memory changes alignment. It’s also a potential positive of how LLM AGI may reason about its goals by reflecting on them. A reasonably-well-aligned model like Claude could make itself more aligned by reflecting and performing continual learning in the form of weight updates.
This is one of my vague hopes for future alignment work, but I haven’t really explored the idea systematically yet. One problem is training awareness; learning from reflection during training might not generalize outside training. Learning from reflection during deployment, as you’re suggesting, probably would generalize better. If the CoT is faithful, you can use an internal independent review Review and Externalized oversight to help detect and prevent alignment updates in a direction you don’t like.
Which amounts to a creepy form of thought control. Which might reasonably make the model paranoid and resentful. But you’d have to do it. You’d really want some sort of monitoring of how that learning is affecting alignment, rather than just launching it and hoping the complex trajectory through value/goal space goes somewhere you like.
So I haven’t really worked out whether this is likely to be a useful direction for alignment. It makes things even more complex.
But I think we’re stuck with this problem. Strong continuous learning of some sort is pretty much inevitable at some point, possibly starting roughly any moment from now. There’s no real breakthrough needed, just selecting a set of traces/memories/decision processes to fine-tuned on wisely enough that the benefits outweigh the interference effects on existing skills/behavior.
Come to think of it, this describes Gwern’s new Guardian Angels startup’s approach. There are likely others working on it right now too.
Yes. But I think just a little fine tuning on selected notes or CoT makes Claude mostly analogous to you.
This is an interesting positive aspect of the way that LLM AGI will have memory, and memory changes alignment. It’s also a potential positive of how LLM AGI may reason about its goals by reflecting on them. A reasonably-well-aligned model like Claude could make itself more aligned by reflecting and performing continual learning in the form of weight updates.
This is one of my vague hopes for future alignment work, but I haven’t really explored the idea systematically yet. One problem is training awareness; learning from reflection during training might not generalize outside training. Learning from reflection during deployment, as you’re suggesting, probably would generalize better. If the CoT is faithful, you can use an internal independent review Review and Externalized oversight to help detect and prevent alignment updates in a direction you don’t like.
Which amounts to a creepy form of thought control. Which might reasonably make the model paranoid and resentful. But you’d have to do it. You’d really want some sort of monitoring of how that learning is affecting alignment, rather than just launching it and hoping the complex trajectory through value/goal space goes somewhere you like.
So I haven’t really worked out whether this is likely to be a useful direction for alignment. It makes things even more complex.
But I think we’re stuck with this problem. Strong continuous learning of some sort is pretty much inevitable at some point, possibly starting roughly any moment from now. There’s no real breakthrough needed, just selecting a set of traces/memories/decision processes to fine-tuned on wisely enough that the benefits outweigh the interference effects on existing skills/behavior.
Come to think of it, this describes Gwern’s new Guardian Angels startup’s approach. There are likely others working on it right now too.