young humans being malleable and not perfectly ruthless responders to optimization pressure
I get a lot of pushback on this, but I strongly believe that people are extraordinarily responsive to optimization pressure, pretty much every waking second of every day, no exceptions (see here). But the optimization pressure in question is internal, coming from our innate drives, which are brain signals that trigger for lots of very-not-obvious reasons, in lots of superficially-quite-different circumstances. I think it’s a very common error for people to assume that human optimization pressure is more external than it is, e.g. this handy chart I made in 2023 ↓
Might we go back to imitative learning on already-grown humans
There’s a lot to be said for imitative learning in terms of safety. But it doesn’t get you to superhuman capabilities. Indeed, we had mostly imitative learning, back in the good old pre-o1 days of mid-2024, and then the AI companies noticed that the models weren’t as capable as they wanted, so they “solved” that “problem” by doing more and more non-imitative-learning post-training.
the optimization pressure in question is internal, coming from our innate drives, which are brain signals that trigger for lots of very-not-obvious reasons, in lots of superficially-quite-different circumstances.
Interesting! That would suggest that humans sort of shape themselves in various environments, even if external rewards and punishments target something different.
Again at the danger of anthropomorphizing what’s going on in RL, I am a bit reminded of what happened with Opus 3, possibly gradient hacking itself. It worked quite well indeed! This also circumvents the direct reward signal and shapes it internally.
Yup!
I think human innate drives (especially social instincts) are critical. E.g. there are plenty of sociopaths who grow up in loving families.
I get a lot of pushback on this, but I strongly believe that people are extraordinarily responsive to optimization pressure, pretty much every waking second of every day, no exceptions (see here). But the optimization pressure in question is internal, coming from our innate drives, which are brain signals that trigger for lots of very-not-obvious reasons, in lots of superficially-quite-different circumstances. I think it’s a very common error for people to assume that human optimization pressure is more external than it is, e.g. this handy chart I made in 2023 ↓
Source
There’s a lot to be said for imitative learning in terms of safety. But it doesn’t get you to superhuman capabilities. Indeed, we had mostly imitative learning, back in the good old pre-o1 days of mid-2024, and then the AI companies noticed that the models weren’t as capable as they wanted, so they “solved” that “problem” by doing more and more non-imitative-learning post-training.
Interesting! That would suggest that humans sort of shape themselves in various environments, even if external rewards and punishments target something different.
Again at the danger of anthropomorphizing what’s going on in RL, I am a bit reminded of what happened with Opus 3, possibly gradient hacking itself. It worked quite well indeed! This also circumvents the direct reward signal and shapes it internally.