which of the following is more important to solve for alignment?
having the ability to determine whatsoever whether any given action is dangerous (but this could be extremely expensive per action)
cheaply determining whether a given action is dangerous, given access to an oracle is always correct about whether it’s dangerous (but the oracle can be extremely expensive)
It seems like mostly a type error to ask if an action is dangerous, or the wrong question. Actions are part of a whole life story of behavior of a mind. There are actions that are definitely dangerous on their own (“press the nuke button”), but all the useful good actions are ambiguous with dangerous actions, if you just look at the action on its own. E.g. “get much smarter” is a dangerous action if the surrounding mind is dangerous, and is good if the surrounding mind is good.
It’s really not (but if you won’t think about it more then I guess there’s not much more to say). At least naively, if you rephrase the question to include the full state, then the question is basically “can you tell whether the mind is aligned or not”. Did you mean to ask something equivalent to that?
hmm, it seems the disagreement is deeper than i initially imagined. to clarify, i mean given access to all of the past observations and actions of the agent.
Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.
I don’t understand what is meant by the second one. Do you mean that the oracle could be extremely expensive up-front, but determines the result for any given query very cheaply?
Does the latter one look like an function that response “Safe|Dangerous|Unknown” and is never wrong about Safe or Dangerous judgements, but might rarely fall back to the oracle that always answers Safe|Dangerous? If that’s the case that seems quite valuable even in the absence of the perfect oracle, if you’ve got something which can categorize between “definitely safe” and “might not be safe” and never judges something as “definitely safe” when it’s actually dangerous, and usually judges safe things as safe (i.e. it’s not just a rock with “might be unsafe” written on it), that seems like it gets you most of the value already.
what counts as an action? if you mean tight definitions e.g. every toolcall is an action then I mostly agree with @TsviBT. you can have a dangerous rollout where no individual action passes any danger threshold, and for notkilleveryoneism risks this is plausibly the default, since long-term plans are diffuse. e.g. an oracle on the legality of official acts wouldn’t necessarily have stopped hitler
which of the following is more important to solve for alignment?
having the ability to determine whatsoever whether any given action is dangerous (but this could be extremely expensive per action)
cheaply determining whether a given action is dangerous, given access to an oracle is always correct about whether it’s dangerous (but the oracle can be extremely expensive)
It seems like mostly a type error to ask if an action is dangerous, or the wrong question. Actions are part of a whole life story of behavior of a mind. There are actions that are definitely dangerous on their own (“press the nuke button”), but all the useful good actions are ambiguous with dangerous actions, if you just look at the action on its own. E.g. “get much smarter” is a dangerous action if the surrounding mind is dangerous, and is good if the surrounding mind is good.
this feels unnecessarily pedantic. you can substitute “state and action” everywhere
It’s really not (but if you won’t think about it more then I guess there’s not much more to say). At least naively, if you rephrase the question to include the full state, then the question is basically “can you tell whether the mind is aligned or not”. Did you mean to ask something equivalent to that?
hmm, it seems the disagreement is deeper than i initially imagined. to clarify, i mean given access to all of the past observations and actions of the agent.
Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.
I don’t understand what is meant by the second one. Do you mean that the oracle could be extremely expensive up-front, but determines the result for any given query very cheaply?
let me rephrase.
you always have the ability to tell whether something is dangerous, but this ability is very expensive
you have the ability to cheaply tell whether something is dangerous, but only if you already have the ability to expensively tell if it’s dangerous
The first one seems necessary to build (or verify) the second one.
No, the second one takes the first one for granted. This is asking which of the two parts of the full solution are load bearing.
Does the latter one look like an function that response “Safe|Dangerous|Unknown” and is never wrong about Safe or Dangerous judgements, but might rarely fall back to the oracle that always answers Safe|Dangerous? If that’s the case that seems quite valuable even in the absence of the perfect oracle, if you’ve got something which can categorize between “definitely safe” and “might not be safe” and never judges something as “definitely safe” when it’s actually dangerous, and usually judges safe things as safe (i.e. it’s not just a rock with “might be unsafe” written on it), that seems like it gets you most of the value already.
the latter is safe|dangerous|unknown.
what counts as an action? if you mean tight definitions e.g. every toolcall is an action then I mostly agree with @TsviBT. you can have a dangerous rollout where no individual action passes any danger threshold, and for notkilleveryoneism risks this is plausibly the default, since long-term plans are diffuse. e.g. an oracle on the legality of official acts wouldn’t necessarily have stopped hitler