It seems like mostly a type error to ask if an action is dangerous, or the wrong question. Actions are part of a whole life story of behavior of a mind. There are actions that are definitely dangerous on their own (“press the nuke button”), but all the useful good actions are ambiguous with dangerous actions, if you just look at the action on its own. E.g. “get much smarter” is a dangerous action if the surrounding mind is dangerous, and is good if the surrounding mind is good.
It’s really not (but if you won’t think about it more then I guess there’s not much more to say). At least naively, if you rephrase the question to include the full state, then the question is basically “can you tell whether the mind is aligned or not”. Did you mean to ask something equivalent to that?
hmm, it seems the disagreement is deeper than i initially imagined. to clarify, i mean given access to all of the past observations and actions of the agent.
Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.
It seems like mostly a type error to ask if an action is dangerous, or the wrong question. Actions are part of a whole life story of behavior of a mind. There are actions that are definitely dangerous on their own (“press the nuke button”), but all the useful good actions are ambiguous with dangerous actions, if you just look at the action on its own. E.g. “get much smarter” is a dangerous action if the surrounding mind is dangerous, and is good if the surrounding mind is good.
this feels unnecessarily pedantic. you can substitute “state and action” everywhere
It’s really not (but if you won’t think about it more then I guess there’s not much more to say). At least naively, if you rephrase the question to include the full state, then the question is basically “can you tell whether the mind is aligned or not”. Did you mean to ask something equivalent to that?
hmm, it seems the disagreement is deeper than i initially imagined. to clarify, i mean given access to all of the past observations and actions of the agent.
Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.