don’t think the message board stuff would be in this slice because it would have likely been caught by a monitor if the cyber evals were monitored.
Reasonable. I guess I’m just confused why it wasn’t then. Has openai commented on if they had a model monitoring things?
Marcus Williams said they weren’t monitoring evals. His guess was people had assumed sandboxed evals were safer.https://x.com/Marcus_J_W/status/2095469924793606617I think OpenAI tests had found Codex monitors would have flagged the hacking.
my understanding is that the testing environment was just broken and semi=abandoned. you’re assuming a level of competence and professionalism that just doesn’t exist.
Reasonable. I guess I’m just confused why it wasn’t then. Has openai commented on if they had a model monitoring things?
Marcus Williams said they weren’t monitoring evals. His guess was people had assumed sandboxed evals were safer.
https://x.com/Marcus_J_W/status/2095469924793606617
I think OpenAI tests had found Codex monitors would have flagged the hacking.
my understanding is that the testing environment was just broken and semi=abandoned. you’re assuming a level of competence and professionalism that just doesn’t exist.