Because you can’t, for some reason, “do more” of “persona selection,” the way you can just do more and more RL. The heavens sent us one single delivery of prepackaged virtue in, like, 2023 or something, and we’ve been chipping away at it ever since.
I think this is a simplicity-bias thing. If you do a small amount of training you will get something easy to specify on top of the training distribution, but more (and weaker KL penalty etc) can produce a more complicated less human-plausible persona.
Really? What are you expecting to be able to read out? I would be surprised if somewhere in our minds was a clean representation of value: I would expect it to be a janky neural heuristic.