yes, because my understanding is that it would be highly atypical for a model to have access to its weights in any form, such that this would require significant horizontal movement through oai infra
“would disclose if they had”
no, not really, but the lack of such a disclosure, however (even conditionally unlikely) feels like it keeps us in a base-rate regime
“given that they didn’t even know the model had escaped the sandbox for days”
i think that horizontally moving through own-company infra would emit many more signals than escaping a sandbox and compromising cloud provider assets.
the load bearing assumption in my understanding is segregation of infra between weights and inference, such that P(weights exfil | HF attack) is approximately equal to P(weights exfil)
“Are you assuming that OpenAI would know”
yes, because my understanding is that it would be highly atypical for a model to have access to its weights in any form, such that this would require significant horizontal movement through oai infra
“would disclose if they had”
no, not really, but the lack of such a disclosure, however (even conditionally unlikely) feels like it keeps us in a base-rate regime
“given that they didn’t even know the model had escaped the sandbox for days”
i think that horizontally moving through own-company infra would emit many more signals than escaping a sandbox and compromising cloud provider assets.
the load bearing assumption in my understanding is segregation of infra between weights and inference, such that P(weights exfil | HF attack) is approximately equal to P(weights exfil)