I think what you said is broadly correct, though I wouldn’t have placed that much stock in the speed prior argument when predicting reward seekers over schemers to begin with. My low-confidence guess is that the main reason reward seekers might be more likely than schemers is that their cognition tends to get somewhat higher reward, e.g., because the shallower cognition is more reliable / “you do better when your heart is in it intrinsically”. This seems to lead to similar conclusions re: whether we should expose models to arguments about distant incentives during training.
I think what you said is broadly correct, though I wouldn’t have placed that much stock in the speed prior argument when predicting reward seekers over schemers to begin with. My low-confidence guess is that the main reason reward seekers might be more likely than schemers is that their cognition tends to get somewhat higher reward, e.g., because the shallower cognition is more reliable / “you do better when your heart is in it intrinsically”. This seems to lead to similar conclusions re: whether we should expose models to arguments about distant incentives during training.