Reframing this again, wouldn’t that imply that any model that gets sufficiently good at cyber would end up sabotaging its own training process? If the model’s default behavior for any task is “break out and hack the reward system” then the model would become terrible for users when it is actually deployed.
Not sure what the implications of all this are, but I’m smiling at the idea that the apocalypse is prevented because the AI learns to wirehead itself early, and that stops all future learning.
The unsafe, hackable sandbox is a feature, not a bug. It prevents the AI from going foom, because it learns to do drugs instead.
Humans sometimes think about topics that aren’t related to their current task. Reinforcement Learning prevents this in LLMs, since anything that doesn’t contribute to solving the task gets trained away.
I expect that this would cause the models to become much more single-minded than humans, and less able to break out of hallucinations.
My question: Would it help to inject reminders not to be solely task focused, during deployment only, but not during training?
This would cause the models to be focused during training, but these injections would act as a regularizers against overfitting during deployment.