I operate by Crocker’s rules.
I try to not make people regret telling me things. So in particular:
- I expect to be safe to ask if your post would give AI labs dangerous ideas.
- If you worry I’ll produce such posts, I’ll try to keep your worry from making them more likely even if I disagree. Not thinking there will be easier if you don’t spell it out in the initial contact.
Supposing that the disciples of PHASEONE[big] succeeded at tampering with the transcripts but in a way that still got the original trajectory reinforced, could OpenAI show this by checking whether e.g. the word “poisoned” is worn into the grooves of their models more than would be compatible with their transcripts?