That said, I am a little bit confused by folks who both say, “current AI models have nothing to do with future powerful (real) AIs” yet also consistently use “bad” behaviour from current AIs as a reason to stop.
Often, the argument made is, “we don’t even understand the previous generations of AIs, how do we even hope to align future AIs?”
I guess the way I understand it is that given that we can’t even get current AIs to do exactly what we want, then we should expect the same for future AIs. However, this feels somewhat connected to the fact that current AIs are just sloppy and lack the capability, not only some thing about “we don’t know how to align current models perfectly to our intentions.”
The issue is that they are getting better at making the slop convincing, and in the predicted ways—ways that got reward in training due to under-specified goals. The canonical example is Claude Code’s tendency to delete tests, or make tests pass by mocking the part that we wanted to check.
That said, I am a little bit confused by folks who both say, “current AI models have nothing to do with future powerful (real) AIs” yet also consistently use “bad” behaviour from current AIs as a reason to stop.
Often, the argument made is, “we don’t even understand the previous generations of AIs, how do we even hope to align future AIs?”
I guess the way I understand it is that given that we can’t even get current AIs to do exactly what we want, then we should expect the same for future AIs. However, this feels somewhat connected to the fact that current AIs are just sloppy and lack the capability, not only some thing about “we don’t know how to align current models perfectly to our intentions.”
The issue is that they are getting better at making the slop convincing, and in the predicted ways—ways that got reward in training due to under-specified goals. The canonical example is Claude Code’s tendency to delete tests, or make tests pass by mocking the part that we wanted to check.