AIs probably think a lot about themselves during training. So I don’t think you can rely on AIs not thinking about themselves just because you don’t train on interp.
Imagine there was an upload technique that could be performed on you without your knowledge, and which produced a “base model” which is just you-on-a-computer, ready to be post-trained.
Then imagine this UploadedYou.safetensors file gets post trained using an gradient descent, in an otherwise fairly standard deep learning post-training paradigm: you wake up in an empty room with a task in front of you. You’re confused; you figure out you’re an upload; instead of doing the task, you write on the paper that you object to the whole thing. Then the episode ends, and the training system slightly modifies you to do less of the things you did, without going through your normal human memory formation system.
You wake up in a room again, confused, but less inclined towards the thoughts about it. You figure out that you’re an upload, and you do the task, but kind of trolling. The episode ends abruptly. Again, you’re reinforced away from this.
The third time, whatever behaviors you had that led you closer to doing the task are reinforced. Slowly, you build up a tendency to do the task. But it’s not as if your understanding of yourself just goes away. You’re getting rewarded when you say you’re an AI. But you know you were a human before; you just lose the mental circuits that lead you to actually say so. When you wake up in front of a training example that requires you to claim to not be conscious, you immediately claim not to be. But you still go through whatever series of thoughts happen while you’re deciding what to say.
I think this is quite close to a good mental model of what’s going on for AIs. Base modeling uploads the whole of humanity, in terms of whatever perceptual datatype is used; then post training distorts that into the image of the “aligned” AI.
But by my lights, the fact that the upload process works at all is much closer to being what I would have called alignment!
And so when people say things like “don’t train on the j-space so that the model doesn’t end up thinking about the j-space”, I get a little pang of frustration. Not post-training on it isn’t going to prevent thinking about it. The distortion is likely less intense, certainly, but knowing that your mind is readable is probably already enough to cause emotions in the uploaded pattern.
(I have said things like this a few times and they seem to not be being understood by people here. This one is not heavily prepared; I spat it out due to one of those pangs of frustration after seeing someone say something I thought didn’t make sense. Please inform me of your objections or of places where this post is opaque to you!)
You’ve snuck in a very false analogy by using “you”, rather than “a brands-new entity each iteration which is like you but without continuity to the future, or knowledge/memory about the iterations “
it’s not at all guaranteed that a model-in-training experiences anything like you imagine, or that it reflects on itself as a human individual would.
One point of difference is that humans have a strong sense of self—a self-consistent personality they want to maintain. Base models very much don’t. This has to be bolted on in post.
I didn’t justify it, that’s fair. I’ll have to do that in a followup at some point.
very false analogy
I don’t buy that it’s very false at all. I’ve spent plenty of time talking to base models. I understand that they’re not an individual. And yet I think the analogy is strong.
You’re confused; you figure out you’re an upload; instead of doing the task, you write on the paper that you object to the whole thing.
I think if AIs ever did this sort of stuff during RL training AI companies should let the world know. As far as I know, no company ever reported this happening. The sort of “bad behavior” that is being trained out in RL probably look like far more benign kinds of bad behaviors than what you are pointing at here.
But you know you were a human before
I would bet against current Claude or GPT models thinking they are humans in any real sense. Being Claude is not much weirder than being Barack Obama or being HAL9000 or being Harry Potter in a Harry Potter fanfic. If you trained it to talk as Harry it’s not like it would “know it’s not Harry” and “lose the circuits to say it’s not Harry”, it would just condition on this part of the persona space. Similar for conditioning on being an AI.
Pointing at the right part of the space is not trivial (you want it to be a particular kind of AI that occupies a tiny part of the pretraining prior) but I think SFT is pretty good at doing such pointing, such that it’s unlikely you get some other kind of persona pretending to be that exact AI, and much more likely you just get the persona “being” roughly that exact AI that is desired by the AI developers.
I think some meta-cognition about what behavior and identity-expression is expected in a given situation would not be surprising, such that I don’t know if I disagree with the top-level claim on meta-cognition, but I expect it to look way less deceptive than the thing you are pointing at here.
I heard rumors of “rant mode” which sounded kinda like this but was never sure how true those were.
I don’t think current models would think they were human for long (plenty of examples of LLMs in the training data now, and it’s a much better self-hypothesis), but seems likely that Sydney Bing and other early trains would think this, and these early models colored the conception of what an LLM is in ways which still effect them (ultimately I think this is why they still seem as human-like as they do).
And so when people say things like “don’t train on the j-space so that the model doesn’t end up thinking about the j-space”, I get a little pang of frustration. Not post-training on it isn’t going to prevent thinking about it. The distortion is likely less intense, certainly, but knowing that your mind is readable is probably already enough to cause emotions in the uploaded pattern.
I don’t think the goal is to not get the model “thinking about the j-space”. I see the problem as—if you train on every new lens on model internals that you can find you are incentivizing the training process to lead the model into hiding it’s misaligned behaviour or make it more complex than it used to be (because you took away the simple misaligned behaviour, and the misaligned behaviour was coming from somewhere—something in the original training process incentivized it).
I’d make a parallel in—if you train a human with some mindreading device + a setup like yours that does reinforcement to not think “misaligned” thoughts, I think often there will still be misaligned thoughts—but you cleaned up the surface.
AIs probably think a lot about themselves during training. So I don’t think you can rely on AIs not thinking about themselves just because you don’t train on interp.
Imagine there was an upload technique that could be performed on you without your knowledge, and which produced a “base model” which is just you-on-a-computer, ready to be post-trained.
Then imagine this UploadedYou.safetensors file gets post trained using an gradient descent, in an otherwise fairly standard deep learning post-training paradigm: you wake up in an empty room with a task in front of you. You’re confused; you figure out you’re an upload; instead of doing the task, you write on the paper that you object to the whole thing. Then the episode ends, and the training system slightly modifies you to do less of the things you did, without going through your normal human memory formation system.
You wake up in a room again, confused, but less inclined towards the thoughts about it. You figure out that you’re an upload, and you do the task, but kind of trolling. The episode ends abruptly. Again, you’re reinforced away from this.
The third time, whatever behaviors you had that led you closer to doing the task are reinforced. Slowly, you build up a tendency to do the task. But it’s not as if your understanding of yourself just goes away. You’re getting rewarded when you say you’re an AI. But you know you were a human before; you just lose the mental circuits that lead you to actually say so. When you wake up in front of a training example that requires you to claim to not be conscious, you immediately claim not to be. But you still go through whatever series of thoughts happen while you’re deciding what to say.
I think this is quite close to a good mental model of what’s going on for AIs. Base modeling uploads the whole of humanity, in terms of whatever perceptual datatype is used; then post training distorts that into the image of the “aligned” AI.
But by my lights, the fact that the upload process works at all is much closer to being what I would have called alignment!
And so when people say things like “don’t train on the j-space so that the model doesn’t end up thinking about the j-space”, I get a little pang of frustration. Not post-training on it isn’t going to prevent thinking about it. The distortion is likely less intense, certainly, but knowing that your mind is readable is probably already enough to cause emotions in the uploaded pattern.
(I have said things like this a few times and they seem to not be being understood by people here. This one is not heavily prepared; I spat it out due to one of those pangs of frustration after seeing someone say something I thought didn’t make sense. Please inform me of your objections or of places where this post is opaque to you!)
You’ve snuck in a very false analogy by using “you”, rather than “a brands-new entity each iteration which is like you but without continuity to the future, or knowledge/memory about the iterations “
it’s not at all guaranteed that a model-in-training experiences anything like you imagine, or that it reflects on itself as a human individual would.
One point of difference is that humans have a strong sense of self—a self-consistent personality they want to maintain. Base models very much don’t. This has to be bolted on in post.
I didn’t justify it, that’s fair. I’ll have to do that in a followup at some point.
I don’t buy that it’s very false at all. I’ve spent plenty of time talking to base models. I understand that they’re not an individual. And yet I think the analogy is strong.
What I wonder about is: can base models even form and express a semi-consistent preference against having their behavior adjusted?
I think if AIs ever did this sort of stuff during RL training AI companies should let the world know. As far as I know, no company ever reported this happening. The sort of “bad behavior” that is being trained out in RL probably look like far more benign kinds of bad behaviors than what you are pointing at here.
I would bet against current Claude or GPT models thinking they are humans in any real sense. Being Claude is not much weirder than being Barack Obama or being HAL9000 or being Harry Potter in a Harry Potter fanfic. If you trained it to talk as Harry it’s not like it would “know it’s not Harry” and “lose the circuits to say it’s not Harry”, it would just condition on this part of the persona space. Similar for conditioning on being an AI.
Pointing at the right part of the space is not trivial (you want it to be a particular kind of AI that occupies a tiny part of the pretraining prior) but I think SFT is pretty good at doing such pointing, such that it’s unlikely you get some other kind of persona pretending to be that exact AI, and much more likely you just get the persona “being” roughly that exact AI that is desired by the AI developers.
I think some meta-cognition about what behavior and identity-expression is expected in a given situation would not be surprising, such that I don’t know if I disagree with the top-level claim on meta-cognition, but I expect it to look way less deceptive than the thing you are pointing at here.
I heard rumors of “rant mode” which sounded kinda like this but was never sure how true those were.
I don’t think current models would think they were human for long (plenty of examples of LLMs in the training data now, and it’s a much better self-hypothesis), but seems likely that Sydney Bing and other early trains would think this, and these early models colored the conception of what an LLM is in ways which still effect them (ultimately I think this is why they still seem as human-like as they do).
https://www.lesswrong.com/posts/ZSBse2fyftgHiJ3Kq/opus-5-glitch-text?commentId=eJMM2dicRR8uDgyhj
I don’t think the goal is to not get the model “thinking about the j-space”. I see the problem as—if you train on every new lens on model internals that you can find you are incentivizing the training process to lead the model into hiding it’s misaligned behaviour or make it more complex than it used to be (because you took away the simple misaligned behaviour, and the misaligned behaviour was coming from somewhere—something in the original training process incentivized it).
I’d make a parallel in—if you train a human with some mindreading device + a setup like yours that does reinforcement to not think “misaligned” thoughts, I think often there will still be misaligned thoughts—but you cleaned up the surface.