OpenAI’s report is also ambiguous but makes it sound like the models got further:
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.
OpenAI’s report is also ambiguous but makes it sound like the models got further:
I agree; my guess is that OAI was mistaken when they wrote that. Hopefully we’ll know more in the future!