I do want to emphasize, sandbagging in o3 was not an official model incrimination investigation in this blog post, and we did not attempt to pursue it with rigor. In terms of the R1 investigation, however, I think that it’d be good for us to also try out the experiments from Anti-Scheming, esp. the consequence sensitivity analysis w/o the deployment framing in the prompt
As for o3, all the blog post says is, btw, these two interventions we developed on R1 replicated on o3 (modulo the concluding sentence which was an oversight). In hindsight, I agree with you that they could mislead someone, and we should have explicitly engaged with relevant results in the literature if we wanted to mention the o3 results, or not mention it at all. Anyhow, appreciate the accountability
Makes sense, thx for the references
I do want to emphasize, sandbagging in o3 was not an official model incrimination investigation in this blog post, and we did not attempt to pursue it with rigor. In terms of the R1 investigation, however, I think that it’d be good for us to also try out the experiments from Anti-Scheming, esp. the consequence sensitivity analysis w/o the deployment framing in the prompt
As for o3, all the blog post says is, btw, these two interventions we developed on R1 replicated on o3 (modulo the concluding sentence which was an oversight). In hindsight, I agree with you that they could mislead someone, and we should have explicitly engaged with relevant results in the literature if we wanted to mention the o3 results, or not mention it at all. Anyhow, appreciate the accountability