1) It doesn’t matter if we don’t think it’s likely. It needs to be demonstrated to a very high confidence level. This is an entirely reasonable ask. 2) I already said there were calls for transparency; the one you reference did not make this demand.
I’m not sure what your ask is, beyond releasing the agent traces, and maybe some additional logs. Is there some other way you propose OA prove that the model didn’t self-exfiltrate? What would a ‘very high confidence level’ entail in addition to this?
You’re eggregiously missing my points:
1) It doesn’t matter if we don’t think it’s likely. It needs to be demonstrated to a very high confidence level. This is an entirely reasonable ask.
2) I already said there were calls for transparency; the one you reference did not make this demand.
I’m not sure what your ask is, beyond releasing the agent traces, and maybe some additional logs. Is there some other way you propose OA prove that the model didn’t self-exfiltrate? What would a ‘very high confidence level’ entail in addition to this?