I’m not sure what your ask is, beyond releasing the agent traces, and maybe some additional logs. Is there some other way you propose OA prove that the model didn’t self-exfiltrate? What would a ‘very high confidence level’ entail in addition to this?
I’m not sure what your ask is, beyond releasing the agent traces, and maybe some additional logs. Is there some other way you propose OA prove that the model didn’t self-exfiltrate? What would a ‘very high confidence level’ entail in addition to this?