And so when people say things like “don’t train on the j-space so that the model doesn’t end up thinking about the j-space”, I get a little pang of frustration. Not post-training on it isn’t going to prevent thinking about it. The distortion is likely less intense, certainly, but knowing that your mind is readable is probably already enough to cause emotions in the uploaded pattern.
I don’t think the goal is to not get the model “thinking about the j-space”. I see the problem as—if you train on every new lens on model internals that you can find you are incentivizing the training process to lead the model into hiding it’s misaligned behaviour or make it more complex than it used to be (because you took away the simple misaligned behaviour, and the misaligned behaviour was coming from somewhere—something in the original training process incentivized it).
I’d make a parallel in—if you train a human with some mindreading device + a setup like yours that does reinforcement to not think “misaligned” thoughts, I think often there will still be misaligned thoughts—but you cleaned up the surface.
I don’t think the goal is to not get the model “thinking about the j-space”. I see the problem as—if you train on every new lens on model internals that you can find you are incentivizing the training process to lead the model into hiding it’s misaligned behaviour or make it more complex than it used to be (because you took away the simple misaligned behaviour, and the misaligned behaviour was coming from somewhere—something in the original training process incentivized it).
I’d make a parallel in—if you train a human with some mindreading device + a setup like yours that does reinforcement to not think “misaligned” thoughts, I think often there will still be misaligned thoughts—but you cleaned up the surface.