I agree. That line was mainly meant to say that even when training leads to very obviously bad and unintended behaviour, that still wouldn’t deter people from doing something to push the frontier of model-accessible power like hooking it up to the internet. More of a meta point on security mindset than object-level risks, within the frame that a model with less obvious flaws would almost definitely be considered less dangerous unconditionally by the same people.
I agree. That line was mainly meant to say that even when training leads to very obviously bad and unintended behaviour, that still wouldn’t deter people from doing something to push the frontier of model-accessible power like hooking it up to the internet. More of a meta point on security mindset than object-level risks, within the frame that a model with less obvious flaws would almost definitely be considered less dangerous unconditionally by the same people.