I would still say this property is not quite necessary when we get strong superintelligence, although it could be the best design choice in practice. The job of preventing people from developing/deploying dangerous AIs doesn’t need to be done by the model’s guardrails; for example we could use surveillance like we do with nuclear weapons.
The requirements for a sufficiently aligned model will depend on what else is happening in society. E.g. if we give everyone weights access to the ASI and don’t regulate AI technology at all, then the AI has to resist being fine-tuned to do dangerous ML research, which is probably infeasible.
I would still say this property is not quite necessary when we get strong superintelligence, although it could be the best design choice in practice. The job of preventing people from developing/deploying dangerous AIs doesn’t need to be done by the model’s guardrails; for example we could use surveillance like we do with nuclear weapons.
The requirements for a sufficiently aligned model will depend on what else is happening in society. E.g. if we give everyone weights access to the ASI and don’t regulate AI technology at all, then the AI has to resist being fine-tuned to do dangerous ML research, which is probably infeasible.