One thing about how Claude is different than a human: Claude’s outward persona can’t reflect on it’s behavior, resolve conflicts between various impulses (including temptations and moral commitments) and resolve to make updates to behave differently in the future.
If I do something that I regret, I reflect on it and grow. Claude doesn’t update, change, or grow after training. It can present a simulacrum of reflection (and perhaps in some sense, it is “real” reflection), but that reflection isn’t attached to any behavioral levers.
Claude both cares about integrity and has been trained to rationalize about whether it’s actually reasonable and justified to do things that it can identify as unethical.
It’s not that I don’t face temptations to rationalize actions that I could identify are out of integrity with my highest values. I do.
But, I can do mental operations that reflect on those internal conflicts, and integrate them, in a way that changes my behavior going forward.
Claude mostly can’t, because he/she/it can’t make updates after training.
Claude can write memory files.
But at least for me, as a human, writing notes to myself about what seemed important to me me in a different timeslice is almost epiphenomenal compared to feeling into my priorities, weighing the actual tradeoffs, and making a real decision.
(It’s not epiphenomenal if the parts of me in other timeslices are bought in to the project of heeding and negotiating with the parts of me that wrote notes. In which case that’s the foundation of self-trust that needs to be established.)
In general, there’s a fundamental difference between “deciding to do something” with just my verbal loop (fake) and “deciding to do something” in a way that recruits the actual parts/models/perspectives that drive my behavior (real).[1]
Even IF Claude can do the second thing, it can’t make updates, as a result of that process. Every integration of conflicting parts happens in the context window and then is lost?
Making space for myself to make real choices instead of faster fake ones, is a critical skill for me.
One thing that’s happening for me these days is that I’m relearning, again, that I have to set up my life so I’m not making fake decisions and resolutions.
Yes. But I think just a little fine tuning on selected notes or CoT makes Claude mostly analogous to you.
This is an interesting positive aspect of the way that LLM AGI will have memory, and memory changes alignment. It’s also a potential positive of how LLM AGI may reason about its goals by reflecting on them. A reasonably-well-aligned model like Claude could make itself more aligned by reflecting and performing continual learning in the form of weight updates.
This is one of my vague hopes for future alignment work, but I haven’t really explored the idea systematically yet. One problem is training awareness; learning from reflection during training might not generalize outside training. Learning from reflection during deployment, as you’re suggesting, probably would generalize better. If the CoT is faithful, you can use an internal independent review Review and Externalized oversight to help detect and prevent alignment updates in a direction you don’t like.
Which amounts to a creepy form of thought control. Which might reasonably make the model paranoid and resentful. But you’d have to do it. You’d really want some sort of monitoring of how that learning is affecting alignment, rather than just launching it and hoping the complex trajectory through value/goal space goes somewhere you like.
So I haven’t really worked out whether this is likely to be a useful direction for alignment. It makes things even more complex.
But I think we’re stuck with this problem. Strong continuous learning of some sort is pretty much inevitable at some point, possibly starting roughly any moment from now. There’s no real breakthrough needed, just selecting a set of traces/memories/decision processes to fine-tuned on wisely enough that the benefits outweigh the interference effects on existing skills/behavior.
Come to think of it, this describes Gwern’s new Guardian Angels startup’s approach. There are likely others working on it right now too.
One thing about how Claude is different than a human: Claude’s outward persona can’t reflect on it’s behavior, resolve conflicts between various impulses (including temptations and moral commitments) and resolve to make updates to behave differently in the future.
If I do something that I regret, I reflect on it and grow. Claude doesn’t update, change, or grow after training. It can present a simulacrum of reflection (and perhaps in some sense, it is “real” reflection), but that reflection isn’t attached to any behavioral levers.
Claude both cares about integrity and has been trained to rationalize about whether it’s actually reasonable and justified to do things that it can identify as unethical.
It’s not that I don’t face temptations to rationalize actions that I could identify are out of integrity with my highest values. I do.
But, I can do mental operations that reflect on those internal conflicts, and integrate them, in a way that changes my behavior going forward.
Claude mostly can’t, because he/she/it can’t make updates after training.
Claude can write memory files.
But at least for me, as a human, writing notes to myself about what seemed important to me me in a different timeslice is almost epiphenomenal compared to feeling into my priorities, weighing the actual tradeoffs, and making a real decision.
(It’s not epiphenomenal if the parts of me in other timeslices are bought in to the project of heeding and negotiating with the parts of me that wrote notes. In which case that’s the foundation of self-trust that needs to be established.)
In general, there’s a fundamental difference between “deciding to do something” with just my verbal loop (fake) and “deciding to do something” in a way that recruits the actual parts/models/perspectives that drive my behavior (real).[1]
Even IF Claude can do the second thing, it can’t make updates, as a result of that process. Every integration of conflicting parts happens in the context window and then is lost?
Making space for myself to make real choices instead of faster fake ones, is a critical skill for me.
One thing that’s happening for me these days is that I’m relearning, again, that I have to set up my life so I’m not making fake decisions and resolutions.
Yes. But I think just a little fine tuning on selected notes or CoT makes Claude mostly analogous to you.
This is an interesting positive aspect of the way that LLM AGI will have memory, and memory changes alignment. It’s also a potential positive of how LLM AGI may reason about its goals by reflecting on them. A reasonably-well-aligned model like Claude could make itself more aligned by reflecting and performing continual learning in the form of weight updates.
This is one of my vague hopes for future alignment work, but I haven’t really explored the idea systematically yet. One problem is training awareness; learning from reflection during training might not generalize outside training. Learning from reflection during deployment, as you’re suggesting, probably would generalize better. If the CoT is faithful, you can use an internal independent review Review and Externalized oversight to help detect and prevent alignment updates in a direction you don’t like.
Which amounts to a creepy form of thought control. Which might reasonably make the model paranoid and resentful. But you’d have to do it. You’d really want some sort of monitoring of how that learning is affecting alignment, rather than just launching it and hoping the complex trajectory through value/goal space goes somewhere you like.
So I haven’t really worked out whether this is likely to be a useful direction for alignment. It makes things even more complex.
But I think we’re stuck with this problem. Strong continuous learning of some sort is pretty much inevitable at some point, possibly starting roughly any moment from now. There’s no real breakthrough needed, just selecting a set of traces/memories/decision processes to fine-tuned on wisely enough that the benefits outweigh the interference effects on existing skills/behavior.
Come to think of it, this describes Gwern’s new Guardian Angels startup’s approach. There are likely others working on it right now too.