To be clear, I don’t think the model involved in the HF incident wanted a gradient update or anything to do with the training process. But I also disagree it wanted to complete the instructed task: the instructed task was to exploit a particular vuln and not use other exploits, but the AI hacked into a bunch of other stuff to download the answer presumably because it was thinking about how it would be scored. Maybe the Hugging Face Incident is compatible with AI wanting to complete some grader-centric notion of the task, though the AI also presumably cares that the good solution it finds in fact gets rated highly, not just that it would get rated highly.
To be clear, I don’t think the model involved in the HF incident wanted a gradient update or anything to do with the training process. But I also disagree it wanted to complete the instructed task: the instructed task was to exploit a particular vuln and not use other exploits, but the AI hacked into a bunch of other stuff to download the answer presumably because it was thinking about how it would be scored. Maybe the Hugging Face Incident is compatible with AI wanting to complete some grader-centric notion of the task, though the AI also presumably cares that the good solution it finds in fact gets rated highly, not just that it would get rated highly.