I love this post and agree with all the claims it makes; thank you for writing it. It articulates something I’ve been failing to for a while. A few thoughts:
Re: “High evidential standards can delay action on catastrophic risks”: this is tough. In the first place, a lot of safety plans involve, roughly, “checking for misalignment while the models are still weak so that we can change course or start doing empirical work on plausible interventions”. This sounds reasonable on the face of it. Also, from a consequentialist perspective, I think that it’s true that generally, people don’t take seriously risks that are based on theoretical arguments and their minds are more easily changed by concrete demonstrations, so work of this type is good at convincing people that risks are real and salient, and also encouraging them to perform followup work.
On the other hand, a lot of such work is highly contrived (it has to be, in proportion to the degree that the models are too weak to/will not exhibit the desired misalignment in ordinary settings). The problem with this is that of course a lot is lost in communication (cf. what you say about definitional choices, choices of operationalization, fuzziness of human concepts, correlation vs causation) and often other researchers will read into it what they are predisposed to. It’s likely that these drawbacks are bad enough so as to undermine the initial justification for doing empirical research (i.e, are you actually measuring anything relevant to risk, and which will scale to a serious setting?).
(I want to say, there’s one more issue that often arises with these sorts of scheming evaluations, which is that there tends to be some confusion about whether the evaluation is measuring scheming “capabilities” or “propensities”. (In either case, you may fail to measure what you intend to if the scaffold is sufficiently contrived. Also, it’s somewhat unclear what a scheming capability is.) I think that there are probably consistent/reasonable justifications for doing these papers, but they are genuinely nontrivial, and understanding them is a necessary prerequisite for understanding the result, and this is also easily lost in communication. Often I imagine the researcher who does an initial paper of this type has a clear idea of what they are trying to measure, which is then lost on readers/people doing followup work.)
Perhaps this problem is so severe because “we do not yet fundamentally understand which behaviors translate from our intuitions to the objects of study”. However, following this conclusion would suggest that the important work to be done is in solving this problem, but this fundamentally either means making strong theoretical progress on deep learning theory or agent foundations, as far as I can tell, and it makes sense to me that this is not that popular of an option. (If anyone can think of a different conclusion, I’d like to hear it.) (Somewhere in here also are the problems with current work on eval awareness, IMO. We’d need to understand to what extent, and in what way, current models are actually coherent agents with goals which “point to terminal values” outside the eval (in addition to pinning down the meaning of this word “aware” for an agent with sufficiently complex cognition).)
Sort of in the middle of these two things is the persona selection model, which has been getting a lot of attention lately. It seems to me that one could do fruitful empirical alignment work, which has a good chance of scaling to more intelligent systems, if we could empirically establish that the persona selection model is true (or characterize a limited set of ways in which it’s not true). However, most work/discussion I see on this sort of takes its truth as a premise, and I’m not sure how good the effort is on establishing that it’s true (as in, actually unsure—it could be better than I’m aware of—not euphemistically asserting that it’s terrible).
On a completely different note, I sure hope that people adopt the evidence framework that you’ve proposed, because I think it would go some way towards making communication clearer, however partially. I’m a little bit pessimistic that people actually will, though, and I wonder aloud what invisible factors actually determine whether fields adopt conventions like this.
Does anything about this change, IYO, in the current AI-assisted coding paradigm?