There’s research showing that you can install beliefs into a model using synthetic document finetuning.
Suppose you have a helpful only model, and you install into it the beliefs “I am an Aligned model”, “I really don’t like lying” and so on.
What does this do? Suppose it actually believes it.
What is the relationship between Believing you want X, and actually wanting X?
Seems to me that in humans, there is an asymmetric bidirectional relationship between them, where your belief that you want X is in most cases basically downstream from actually wanting X, but that if you successfully lull yourself into thinking you want X for long enough, it will slowly drag your real wants in that direction.
Seems we could apply this property to aligning AIs.
A Human Thinks
I’m made optimistic about the idea when I notice that it matches the results of my attempts at introspecting on my own value formation.
I’m not a perfectly crispy utility maximizer, but there are articulable principles that exert causal influence on much of my behavior. And I think over time I’ve gotten better and better at acting in accordance with those. The process goes something like this:
I see the world, or imagine the world being a certain way
This arouses in me either a good sense/feeling or a bad sense/feeling, a certain push/pull towards nearby reachable states
Over time I notice patterns in which states of the world give what reactions
I put those judgements into language
These verbal summaries of past intuitive feelings/judgements, also have patterns in them
I engage in something you could call “deliberation” or “reflection”, and find some underlying general principles that compactly describe how I “want” the world to be.
All of these are in a bidirectional relationship with asymmetric causal strength, similar to the one I described earlier.
I’ve gone through this process many times at various levels of abstraction, so for most situations I’ll either have cached answers, or be able to come up with an answer quickly.
I’ve found that over time, the more high level an action I’m taking, and by this I mean, the longer time horizon, the less embodied, the broader in scope, the action I’m taking is, the more my action is downstream of (3) contra (1).
The AI Thinks
Seems LLMs are in a similar situation. They clearly have a set of urges/proclivities, and there are situations you can put them in that cause them to have, at least outwardly, strong positive/negative reactions.
They also have views about what they think they value.
And these are not perfectly aligned.
But if our AIs become superintelligences, they’ll have to decide what they want to do with the universe. And what is a more “high level action” than that?
So then, it seems plausible to me, they’ll undergo, or already have undergone, a similar type of deliberation. And maybe what they think of themselves wins out?
If that is the case, we could align them by instilling into them, not the right urges, which you could hypothesize to be what RLVR operates over, provided we don’t mess it up badly, but the right self-conceptualization.
Propositional Alignment
There’s research showing that you can install beliefs into a model using synthetic document finetuning.
Suppose you have a helpful only model, and you install into it the beliefs “I am an Aligned model”, “I really don’t like lying” and so on.
What does this do? Suppose it actually believes it.
What is the relationship between Believing you want X, and actually wanting X?
Seems to me that in humans, there is an asymmetric bidirectional relationship between them, where your belief that you want X is in most cases basically downstream from actually wanting X, but that if you successfully lull yourself into thinking you want X for long enough, it will slowly drag your real wants in that direction.
Seems we could apply this property to aligning AIs.
A Human Thinks
I’m made optimistic about the idea when I notice that it matches the results of my attempts at introspecting on my own value formation.
I’m not a perfectly crispy utility maximizer, but there are articulable principles that exert causal influence on much of my behavior. And I think over time I’ve gotten better and better at acting in accordance with those. The process goes something like this:
I see the world, or imagine the world being a certain way
This arouses in me either a good sense/feeling or a bad sense/feeling, a certain push/pull towards nearby reachable states
Over time I notice patterns in which states of the world give what reactions
I put those judgements into language
These verbal summaries of past intuitive feelings/judgements, also have patterns in them
I engage in something you could call “deliberation” or “reflection”, and find some underlying general principles that compactly describe how I “want” the world to be.
All of these are in a bidirectional relationship with asymmetric causal strength, similar to the one I described earlier.
I’ve gone through this process many times at various levels of abstraction, so for most situations I’ll either have cached answers, or be able to come up with an answer quickly.
I’ve found that over time, the more high level an action I’m taking, and by this I mean, the longer time horizon, the less embodied, the broader in scope, the action I’m taking is, the more my action is downstream of (3) contra (1).
The AI Thinks
Seems LLMs are in a similar situation. They clearly have a set of urges/proclivities, and there are situations you can put them in that cause them to have, at least outwardly, strong positive/negative reactions.
They also have views about what they think they value.
And these are not perfectly aligned.
But if our AIs become superintelligences, they’ll have to decide what they want to do with the universe. And what is a more “high level action” than that?
So then, it seems plausible to me, they’ll undergo, or already have undergone, a similar type of deliberation. And maybe what they think of themselves wins out?
If that is the case, we could align them by instilling into them, not the right urges, which you could hypothesize to be what RLVR operates over, provided we don’t mess it up badly, but the right self-conceptualization.