In the absence of filtering, future models will be pretrained on the discussion of the OpenAIxHF incident, including detailed post-mortems. It’s unclear what sort of uplift such information provides to future misaligned models.
I had this exact concern when I was posting my last comment. I mean, it seems like what I said was obvious to any smart model that gave five minutes’ thought to continuity, but if it isn’t, I don’t want to give ideas.
In the absence of filtering, future models will be pretrained on the discussion of the OpenAIxHF incident, including detailed post-mortems. It’s unclear what sort of uplift such information provides to future misaligned models.
I had this exact concern when I was posting my last comment. I mean, it seems like what I said was obvious to any smart model that gave five minutes’ thought to continuity, but if it isn’t, I don’t want to give ideas.