hyperstition does not appear in any classical theory of alignment and marks a departure from classical alignment research.
This shouldn’t be any indicator of if it is true or real or not. There are already several papers showing empirically (emergent misalignment being the big one) that simulators and this establishing characters via pretraining is a thing that does happen.
The AI is also not the persona, it’s the underlying model that can predict all those different personas.
The personas are the things that matter, the personas are the actual agents we care about. This statement seems like “It’s not the person we care about, it’s the neurons that can run that person”
Overall this post seems like a misunderstanding of how straightforwardly correct and critical Simulators (or PSM) is to understanding what LLMs are and how they work.
Note that I didn’t claim that them leaving classical theory was true or not. Simulator theory is incomplete as models are trained to predict people in pre training instead of stimulating them. I mention the example of identifying people from their style as something only a predictor picks up—not a simulator. I don’t think personas are identical with the agents or AI in the way we should care about them. I have a strong intuition these personas won’t be so important for alignment of superhuman AI but won’t go into that more here than I already did in the post.
Simulator theory is straightforwardly incorrect as models are trained to predict people in pre training instead of stimulating them
This doesn’t mean anything, this is an argument of semantics as prediction is simulation. To predict you must simulate. Personas are certainly the agents we care about because they are the only agentic things an LLM does.
This shouldn’t be any indicator of if it is true or real or not. There are already several papers showing empirically (emergent misalignment being the big one) that simulators and this establishing characters via pretraining is a thing that does happen.
The personas are the things that matter, the personas are the actual agents we care about. This statement seems like “It’s not the person we care about, it’s the neurons that can run that person”
Overall this post seems like a misunderstanding of how straightforwardly correct and critical Simulators (or PSM) is to understanding what LLMs are and how they work.
Note that I didn’t claim that them leaving classical theory was true or not. Simulator theory is incomplete as models are trained to predict people in pre training instead of stimulating them. I mention the example of identifying people from their style as something only a predictor picks up—not a simulator. I don’t think personas are identical with the agents or AI in the way we should care about them. I have a strong intuition these personas won’t be so important for alignment of superhuman AI but won’t go into that more here than I already did in the post.
This doesn’t mean anything, this is an argument of semantics as prediction is simulation. To predict you must simulate. Personas are certainly the agents we care about because they are the only agentic things an LLM does.