Not sure if I should want you to answer this publicly, but I’m confused how you get strong misaligned results through the OpenAI API? When I try similar fine-tuning experiments through the API I get blocked by moderation checks
OpenAI has given us access to API finetuning with moderation checks disabled, as part of the researcher access program. This is stated in the acknowledgements to the paper. Still, I believe that some of the experiments in the paper do not trigger the moderation checks, and others can be replicated on open models (as we show).
Not sure if I should want you to answer this publicly, but I’m confused how you get strong misaligned results through the OpenAI API? When I try similar fine-tuning experiments through the API I get blocked by moderation checks
OpenAI has given us access to API finetuning with moderation checks disabled, as part of the researcher access program. This is stated in the acknowledgements to the paper. Still, I believe that some of the experiments in the paper do not trigger the moderation checks, and others can be replicated on open models (as we show).