fwiw, my pledge would have forbidden me from working on directly making RLHF work better, or working on HHH, or model spec stuff, etc (and indeed, i have never worked on such things at openai, even though i have had many opportunities to; the closest is my RLHF goodharting work, which intentionally focuses on how to study goodharting rather than how to make RLHF better in general).
fwiw, my pledge would have forbidden me from working on directly making RLHF work better, or working on HHH, or model spec stuff, etc (and indeed, i have never worked on such things at openai, even though i have had many opportunities to; the closest is my RLHF goodharting work, which intentionally focuses on how to study goodharting rather than how to make RLHF better in general).