Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.
Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.