Thanks, that post and the comments mirrors a lot of what I’ve been seeing with Opus 4.8.
I’m probably a bit more pessimistic about the issue being solved any time soon. I don’t see any way to fix this for most kinds of tasks without the RLHF/RLAIF reviewer just being smart enough to catch instances of deception. (Prediction: RLAIF will or already has made the issue worse?) So long as the alignment problem is unsolved, I don’t see how you solve this.
You might be interested in Current AIs seem pretty misaligned to me by ryan greenblatt
Thanks, that post and the comments mirrors a lot of what I’ve been seeing with Opus 4.8.
I’m probably a bit more pessimistic about the issue being solved any time soon. I don’t see any way to fix this for most kinds of tasks without the RLHF/RLAIF reviewer just being smart enough to catch instances of deception. (Prediction: RLAIF will or already has made the issue worse?) So long as the alignment problem is unsolved, I don’t see how you solve this.