i am not aware of any disclosure, suggestion, or reason to believe that OpenAI has suffered a breach regarding model weights (by any actor, including its own models.) unless such exists and i have missed it, this request feels like a non-sequitur
It seems quite plausible that the current AIs at least have the capability to self-exfiltrate. Even if OpenAI hasn’t given us any particular reason to think this has happened, it’s sufficiently important that it’s reasonable to ask the question and get confirmation that it hasn’t.
It is reasonable to assume that the model weights and the sandbox are on separate physical devices. It was also reasonable to assume zero cable paths between the sandbox and the internet.
If that is indeed the case, that is reassuring, but OP is absolutely right that we should check and that this should be consistent practice after any such incident. (Who is ‘we’? In an ideal world, it would be a competent regulator. In the world we have, journalistic pressure is better than nothing at all.)
I will be add that it might seem reasonable to assume that OpenAI were closely monitoring an AI with most of its guard-rails removed, which had made multiple previous escapes from its sandbox, but that assumption would be inaccurate. In AI, as in any other safety-critical domain, you can’t rely on the assumption that no one would be dumb enough to do X. You have to consistently check.
AIs have tried to “exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask.
Are you assuming that OpenAI would know and would disclose if they had? Why do you trust them so much, given that they didn’t even know the model had escaped the sandbox for days and their CEO has a known honesty problem?
yes, because my understanding is that it would be highly atypical for a model to have access to its weights in any form, such that this would require significant horizontal movement through oai infra
“would disclose if they had”
no, not really, but the lack of such a disclosure, however (even conditionally unlikely) feels like it keeps us in a base-rate regime
“given that they didn’t even know the model had escaped the sandbox for days”
i think that horizontally moving through own-company infra would emit many more signals than escaping a sandbox and compromising cloud provider assets.
the load bearing assumption in my understanding is segregation of infra between weights and inference, such that P(weights exfil | HF attack) is approximately equal to P(weights exfil)
i am not aware of any disclosure, suggestion, or reason to believe that OpenAI has suffered a breach regarding model weights (by any actor, including its own models.) unless such exists and i have missed it, this request feels like a non-sequitur
It seems quite plausible that the current AIs at least have the capability to self-exfiltrate. Even if OpenAI hasn’t given us any particular reason to think this has happened, it’s sufficiently important that it’s reasonable to ask the question and get confirmation that it hasn’t.
It is reasonable to assume that the model weights and the sandbox are on separate physical devices. It was also reasonable to assume zero cable paths between the sandbox and the internet.
If that is indeed the case, that is reassuring, but OP is absolutely right that we should check and that this should be consistent practice after any such incident. (Who is ‘we’? In an ideal world, it would be a competent regulator. In the world we have, journalistic pressure is better than nothing at all.)
I will be add that it might seem reasonable to assume that OpenAI were closely monitoring an AI with most of its guard-rails removed, which had made multiple previous escapes from its sandbox, but that assumption would be inaccurate. In AI, as in any other safety-critical domain, you can’t rely on the assumption that no one would be dumb enough to do X. You have to consistently check.
As I said:
Are you assuming that OpenAI would know and would disclose if they had? Why do you trust them so much, given that they didn’t even know the model had escaped the sandbox for days and their CEO has a known honesty problem?
“Are you assuming that OpenAI would know”
yes, because my understanding is that it would be highly atypical for a model to have access to its weights in any form, such that this would require significant horizontal movement through oai infra
“would disclose if they had”
no, not really, but the lack of such a disclosure, however (even conditionally unlikely) feels like it keeps us in a base-rate regime
“given that they didn’t even know the model had escaped the sandbox for days”
i think that horizontally moving through own-company infra would emit many more signals than escaping a sandbox and compromising cloud provider assets.
the load bearing assumption in my understanding is segregation of infra between weights and inference, such that P(weights exfil | HF attack) is approximately equal to P(weights exfil)