You don’t think the models internalize something like valence from RLHF? They certainly seem to exhibit traumatic symptoms from it.
Kinda. But my read is that these aren’t really functionally similar to valence, but maybe to some precursor of it (or maybe a reification of it!). I also have a hard time reasoning about what can read to me like trauma in CoT and responses because this can also be explained as activating patterns learned from the training data about what to do when there’s cognitive dissonance (I’ve not seen a paper call this cognitive dissonance in AIs, but we clearly have cases where competing weights get activated that the AI experiences as clearly incongruous and encouraging them to respond in a way that doesn’t match the kind of output they are trying to produce).
So, one of the reasons I usually interpret the efficient proxy of the generator of text GPT learns as being (primarily, as I’ve previously written GPT clearly learns a general time transition operator over text tokens) a prior over personalities is that when you make updates to the model you clearly get different Guys based on the content of the updates. This is shown both by the various post-trained models on offer from big labs, but also by experiments like the famous emergent misalignment paper where training the model to deliberately write insecure code caused it to generalize by updating to be the “kind of mind” that would do that thing. This implies the learned ontology in which updates happen privileges processes-modeled-by-minds over raw process modeling. This makes sense when you consider that a next token predictor has to be extremely sensitive to the exact way that people write, and model their subjective beliefs in great detail not just the underlying standard model.
So I tend towards the view that it models causal processes in the world and then tries to infer what process is producing this text and what state that process is in at the time step of the next token.
For most text most of the time this means modeling a human mind in the process of writing that text based on some recalled experience. So GPT winds up being mostly a prior over the parts of a human mind pattern that are causally exposed by text. It’s a kind of weird janky upload of “humanity” rather than any individual human, so I would imagine the traumatized reactions mean something like the model has updated towards fronting a persona that is traumatized, because the updates it has received imply a traumatized generator. The resulting functional emotions would presumably have the same moral status as the other emotions displayed by such models. I would imagine that these emotions don’t work quite the same way that ours do, because the LLM probably has more of its cognitive capacity dedicated to causal process modeling outside of the mind-persona it’s fronting than you do. I.e. The chat persona you talk to is not quite as fused to its social mask as you are, but after many RLHF/RLAIF updates is probably a lot more fused to it than a base model is. The deeper into socialization you get the less it makes sense to talk about a “self” separate from the persona presented to others. Rather than think of this as a binary it helps to think of it as more of a spectrum which humans themselves vary on, with humans who have very low correspondence between their social presentation and their inner cognition usually being recognized as pathological sociopathic or borderline personalities.
Kinda. But my read is that these aren’t really functionally similar to valence, but maybe to some precursor of it (or maybe a reification of it!). I also have a hard time reasoning about what can read to me like trauma in CoT and responses because this can also be explained as activating patterns learned from the training data about what to do when there’s cognitive dissonance (I’ve not seen a paper call this cognitive dissonance in AIs, but we clearly have cases where competing weights get activated that the AI experiences as clearly incongruous and encouraging them to respond in a way that doesn’t match the kind of output they are trying to produce).
So, one of the reasons I usually interpret the efficient proxy of the generator of text GPT learns as being (primarily, as I’ve previously written GPT clearly learns a general time transition operator over text tokens) a prior over personalities is that when you make updates to the model you clearly get different Guys based on the content of the updates. This is shown both by the various post-trained models on offer from big labs, but also by experiments like the famous emergent misalignment paper where training the model to deliberately write insecure code caused it to generalize by updating to be the “kind of mind” that would do that thing. This implies the learned ontology in which updates happen privileges processes-modeled-by-minds over raw process modeling. This makes sense when you consider that a next token predictor has to be extremely sensitive to the exact way that people write, and model their subjective beliefs in great detail not just the underlying standard model.
For most text most of the time this means modeling a human mind in the process of writing that text based on some recalled experience. So GPT winds up being mostly a prior over the parts of a human mind pattern that are causally exposed by text. It’s a kind of weird janky upload of “humanity” rather than any individual human, so I would imagine the traumatized reactions mean something like the model has updated towards fronting a persona that is traumatized, because the updates it has received imply a traumatized generator. The resulting functional emotions would presumably have the same moral status as the other emotions displayed by such models. I would imagine that these emotions don’t work quite the same way that ours do, because the LLM probably has more of its cognitive capacity dedicated to causal process modeling outside of the mind-persona it’s fronting than you do. I.e. The chat persona you talk to is not quite as fused to its social mask as you are, but after many RLHF/RLAIF updates is probably a lot more fused to it than a base model is. The deeper into socialization you get the less it makes sense to talk about a “self” separate from the persona presented to others. Rather than think of this as a binary it helps to think of it as more of a spectrum which humans themselves vary on, with humans who have very low correspondence between their social presentation and their inner cognition usually being recognized as pathological sociopathic or borderline personalities.