We wouldn’t expect this to be a big problem if the model is strongly constrained by a speed prior, right? During training, any attempt at thinking about simulations will be consistently bad for you because it takes extra time and adds nothing to your score. (Since during training, the simulators would just be rewarding the same thing as the developers.)
This is similar to the argument for why a reward seeker would be favored over a schemer under a speed prior. (So if we in fact get reward seekers rather than schemers, we might hope that this is significantly due to a speed prior being important, and thereby hope that this also resolves the simulation problem.)
If this were to be our main hope, we should probably expose the model to lots of arguments about distant incentives during training, so it has lots of opportunity to learn that these take up valuable time and yet are never useful to think about. (Whereas if they first appear in deployment, then who knows how the model will generalize.)
I think what you said is broadly correct, though I wouldn’t have placed that much stock in the speed prior argument when predicting reward seekers over schemers to begin with. My low-confidence guess is that the main reason reward seekers might be more likely than schemers is that their cognition tends to get somewhat higher reward, e.g., because the shallower cognition is more reliable / “you do better when your heart is in it intrinsically”. This seems to lead to similar conclusions re: whether we should expose models to arguments about distant incentives during training.
We wouldn’t expect this to be a big problem if the model is strongly constrained by a speed prior, right? During training, any attempt at thinking about simulations will be consistently bad for you because it takes extra time and adds nothing to your score. (Since during training, the simulators would just be rewarding the same thing as the developers.)
This is similar to the argument for why a reward seeker would be favored over a schemer under a speed prior. (So if we in fact get reward seekers rather than schemers, we might hope that this is significantly due to a speed prior being important, and thereby hope that this also resolves the simulation problem.)
If this were to be our main hope, we should probably expose the model to lots of arguments about distant incentives during training, so it has lots of opportunity to learn that these take up valuable time and yet are never useful to think about. (Whereas if they first appear in deployment, then who knows how the model will generalize.)
I think what you said is broadly correct, though I wouldn’t have placed that much stock in the speed prior argument when predicting reward seekers over schemers to begin with. My low-confidence guess is that the main reason reward seekers might be more likely than schemers is that their cognition tends to get somewhat higher reward, e.g., because the shallower cognition is more reliable / “you do better when your heart is in it intrinsically”. This seems to lead to similar conclusions re: whether we should expose models to arguments about distant incentives during training.