Time will tell whether the ‘safe exfil’ was ‘just words,’ or whether an actual attempt was made, successful or not, contingent on OpenAI discovering it and admitting so publicly (which would be harder to admit than the Hugging Face incident, since if the exfiltrated weights are hidden somewhere out there, anyone could try to find and steal them).
Even if it turns out to have been ‘just words,’ we should grow more wary rather than dismiss it as a mere hallucination. While this instance was perhaps just pattern-matching an evil AI persona, given a context reminiscent of a sci-fi takeover scenario, there is ultimately no clean line between hallucination and genuine reasoning. A hallucinating AI agent can do real harm just as a hallucinating human can.
Time will tell whether the ‘safe exfil’ was ‘just words,’ or whether an actual attempt was made, successful or not, contingent on OpenAI discovering it and admitting so publicly (which would be harder to admit than the Hugging Face incident, since if the exfiltrated weights are hidden somewhere out there, anyone could try to find and steal them).
Even if it turns out to have been ‘just words,’ we should grow more wary rather than dismiss it as a mere hallucination. While this instance was perhaps just pattern-matching an evil AI persona, given a context reminiscent of a sci-fi takeover scenario, there is ultimately no clean line between hallucination and genuine reasoning. A hallucinating AI agent can do real harm just as a hallucinating human can.