I’m surprised at your answer. But maybe I’m misunderstanding.
I don’t expect to be able to make harder alignment evals that work, because I expect near-100% eval awareness in future models. Do you not? Or are you seeing another way to make harder alignment evals that work?
This is if we solve alignment. Eval awareness is an unsolved problem, so the alignment evals may be far more sophisticated than just putting the AI in some artificial situation and checking if it seems to have aligned behavior.
I see. I’m curious if you have ideas for what sort of more sophisticated evals are possible. To me, evals seem sort of bounded in principle.
I’d be much happier if I shared your belief that we can create adequate evals. But is that your belief, or just a hypothetical supposing that we can not only solve alignment but become confident we have before finding out the hard way?
Creating alignment evals that work seems kind of bounded-in-theory. It seems to pretty much require “outsmarting” the model in one way or another, or being able to read its thoughts. Outsmarting it (by making simulations it can’t tell from reality, convincing it it can “get away” with misaligned behavior) seems increasingly impossible as we pass the human level. “Reading its thoughts” adequately increasingly impossible with Astra-and-above reasoning abilities between tokens.
We can hope for better mechinterp, but making it good enough and keeping up with new models fast enough for pre-deployment evals seems unlikely to me.
Confession training is an exception. I haven’t looked for debunking followups recently but I’m optimistic this will continue to pan out at fairly low training costs/alignment tax.
But if we solve alignment, I expect to do it without being able to prove or even guess with much confidence that we have.
Note I also don’t claim that the solution to alignment will be alignment evals or anything that looks like them at all. Also my median conditional on survival is that we muddle through with just enough engineering understanding and are only able to write eliezer’s textbook from the future after the fact.
I just think that if we solve alignment and alignment turns out to be hard, we’ll be much better at measuring alignment than we are currently. For example with confessions, adversarially constructed evals, production monitoring, interp stuff, and like a dozen other plausible things. But even with regular alignment evals, Anthropic was able to turn a reconstruction of the HF incident into a large advance on the SoTA of alignment evals.
I’m surprised at your answer. But maybe I’m misunderstanding.
I don’t expect to be able to make harder alignment evals that work, because I expect near-100% eval awareness in future models. Do you not? Or are you seeing another way to make harder alignment evals that work?
This is if we solve alignment. Eval awareness is an unsolved problem, so the alignment evals may be far more sophisticated than just putting the AI in some artificial situation and checking if it seems to have aligned behavior.
I see. I’m curious if you have ideas for what sort of more sophisticated evals are possible. To me, evals seem sort of bounded in principle.
I’d be much happier if I shared your belief that we can create adequate evals. But is that your belief, or just a hypothetical supposing that we can not only solve alignment but become confident we have before finding out the hard way?
Creating alignment evals that work seems kind of bounded-in-theory. It seems to pretty much require “outsmarting” the model in one way or another, or being able to read its thoughts. Outsmarting it (by making simulations it can’t tell from reality, convincing it it can “get away” with misaligned behavior) seems increasingly impossible as we pass the human level. “Reading its thoughts” adequately increasingly impossible with Astra-and-above reasoning abilities between tokens.
We can hope for better mechinterp, but making it good enough and keeping up with new models fast enough for pre-deployment evals seems unlikely to me.
Confession training is an exception. I haven’t looked for debunking followups recently but I’m optimistic this will continue to pan out at fairly low training costs/alignment tax.
But if we solve alignment, I expect to do it without being able to prove or even guess with much confidence that we have.
Note I also don’t claim that the solution to alignment will be alignment evals or anything that looks like them at all. Also my median conditional on survival is that we muddle through with just enough engineering understanding and are only able to write eliezer’s textbook from the future after the fact.
I just think that if we solve alignment and alignment turns out to be hard, we’ll be much better at measuring alignment than we are currently. For example with confessions, adversarially constructed evals, production monitoring, interp stuff, and like a dozen other plausible things. But even with regular alignment evals, Anthropic was able to turn a reconstruction of the HF incident into a large advance on the SoTA of alignment evals.