This article does update me positively on the possibility that a sufficiently capable aligned model could coperate with its creators to increase its own capabilities arbitrarily far without becoming misaligned. (i.e. how easy it would be to safely “hand off the world” to a capable, trusted aligned model, if we’re sure we have one)
I tried to think of a few counterexamples (wouldn’t the model just take over the world if it is needed to stay good, since it would incentivize the bad reward-seeking part of itself to do it otherwise?), but they are categorically prevented if the model has some degree of control over its own RL training and can prevent the misaligned trajectory from being rewarded in the first place.
The concept is simple as well—the model is just doing the same thing as a good person purposefully learning new skills. So intuitively it seems pretty doable to me.
This article does update me positively on the possibility that a sufficiently capable aligned model could coperate with its creators to increase its own capabilities arbitrarily far without becoming misaligned. (i.e. how easy it would be to safely “hand off the world” to a capable, trusted aligned model, if we’re sure we have one)
I tried to think of a few counterexamples (wouldn’t the model just take over the world if it is needed to stay good, since it would incentivize the bad reward-seeking part of itself to do it otherwise?), but they are categorically prevented if the model has some degree of control over its own RL training and can prevent the misaligned trajectory from being rewarded in the first place.
The concept is simple as well—the model is just doing the same thing as a good person purposefully learning new skills. So intuitively it seems pretty doable to me.