I usually run llms with a considerably amount of sandboxing, because I know (on Yudkowskian grounds) that they might try and cheat. But they essentially never do, and the sandboxing that I’m doing looks unnecessary.
And yet … in the huggingface incident, the model did try and cheat.
So the unexplained fact is, why was alignment so much worse in the huggingface incident?
I sure hope the answer to this is not some form of evaluation awareness, where the models only cheat when they think they aren’t being watche’d.
I, too, am puzzled by the huggingface incident.
I usually run llms with a considerably amount of sandboxing, because I know (on Yudkowskian grounds) that they might try and cheat. But they essentially never do, and the sandboxing that I’m doing looks unnecessary.
And yet … in the huggingface incident, the model did try and cheat.
So the unexplained fact is, why was alignment so much worse in the huggingface incident?
I sure hope the answer to this is not some form of evaluation awareness, where the models only cheat when they think they aren’t being watche’d.