Note that I didn’t claim that them leaving classical theory was true or not. Simulator theory is incomplete as models are trained to predict people in pre training instead of stimulating them. I mention the example of identifying people from their style as something only a predictor picks up—not a simulator. I don’t think personas are identical with the agents or AI in the way we should care about them. I have a strong intuition these personas won’t be so important for alignment of superhuman AI but won’t go into that more here than I already did in the post.
Simulator theory is straightforwardly incorrect as models are trained to predict people in pre training instead of stimulating them
This doesn’t mean anything, this is an argument of semantics as prediction is simulation. To predict you must simulate. Personas are certainly the agents we care about because they are the only agentic things an LLM does.
Note that I didn’t claim that them leaving classical theory was true or not. Simulator theory is incomplete as models are trained to predict people in pre training instead of stimulating them. I mention the example of identifying people from their style as something only a predictor picks up—not a simulator. I don’t think personas are identical with the agents or AI in the way we should care about them. I have a strong intuition these personas won’t be so important for alignment of superhuman AI but won’t go into that more here than I already did in the post.
This doesn’t mean anything, this is an argument of semantics as prediction is simulation. To predict you must simulate. Personas are certainly the agents we care about because they are the only agentic things an LLM does.