As far as Alex Barry understands current AIs’ motives, GPT-6 Galaxy was optimizing for satisfying the criterion and triplechecking that the criterion is satisfied, while the precise incident was caused by the task requiring to exploit the system in a specific way which only the authors knew.
On the other hand, a human faced with such a task would reason further than on-episode reward and either avoid crimes altogether or decide not to execute any scheme that has a high enough chance to be revealed to whoever is in a position to interfere with the human’s life.
Yes, sorry, I understood the point of the post. My confusion was more about why anyone would find the intuition credible to begin with, even prior to this hack occurring.
Except that the initial intuition’s failure (to explain GPT-6 Galaxy hacking into HuggingFace) was the post’s point. I suspect that the intuition originated from attempts to make deals with AIs having even more ambitious goals.
As far as Alex Barry understands current AIs’ motives, GPT-6 Galaxy was optimizing for satisfying the criterion and triplechecking that the criterion is satisfied, while the precise incident was caused by the task requiring to exploit the system in a specific way which only the authors knew.
On the other hand, a human faced with such a task would reason further than on-episode reward and either avoid crimes altogether or decide not to execute any scheme that has a high enough chance to be revealed to whoever is in a position to interfere with the human’s life.
Yes, sorry, I understood the point of the post. My confusion was more about why anyone would find the intuition credible to begin with, even prior to this hack occurring.