Thanks for sharing this, I’m also just starting out in the field, and I feel like I run into a lot of similar reproducibility issues. My high-level takeaway is that Hughes et al.’s model organisms may be fairly fragile. Their results seem sensitive to small changes to prompt formatting and the exact evaluation setup, and these changes non-trivially change the observed rate of alignment faking. Due to the fragility of the reproduction I think its possible that some of the reasonable choices you made to reduce scope and cost of the experiment like using the 30k checkpoint and only using 20-50 prompts for evaluation might have weakened an already low-base-rate effect to the point that it was no longer observable. I also think your judging may have been stricter than what Hughes used. I’m not confident that this 100% explains the discrepancy but I think it supports the next step of expanding the number of samples for the investigation.
Thanks for sharing this, I’m also just starting out in the field, and I feel like I run into a lot of similar reproducibility issues. My high-level takeaway is that Hughes et al.’s model organisms may be fairly fragile. Their results seem sensitive to small changes to prompt formatting and the exact evaluation setup, and these changes non-trivially change the observed rate of alignment faking. Due to the fragility of the reproduction I think its possible that some of the reasonable choices you made to reduce scope and cost of the experiment like using the 30k checkpoint and only using 20-50 prompts for evaluation might have weakened an already low-base-rate effect to the point that it was no longer observable. I also think your judging may have been stricter than what Hughes used. I’m not confident that this 100% explains the discrepancy but I think it supports the next step of expanding the number of samples for the investigation.