TsviBT comments on The Shutdown Problem: Incomplete Preferences as a Solution

TsviBT 2 Apr 2024 19:38 UTC
LW: 6 AF: 5
2
AF
IDK if this is a crux for me thinking this is very relevant to stuff on my perspective, but:

The training procedure you propose doesn’t seem to actually incentivize indifference. First, a toy model where I agree it does incentivize that:

On the first time step, the agent gets a choice: choose a number 1--N. If the agent says k, then the agent has nothing at all to do for the first k steps, after which some game G starts. (Each play of G is i.i.d., not related to k.)

So this agent is indeed incentivized to pick k uniformly at random from 1--N. Now consider:

The agent is in a rich world. There are many complex multi-step plans to incentivize agent to learn problem-solving. Each episode, at time N, the agent gets to choose: end now, or play 10 more steps.

Does this incentivize random choice at time N? No. It incentivizes the agent to choose randomly End or Continue at the very beginning of the episode, and then carefully plan and execute behavior that acheives the most reward assuming a run of length N or N+10 respectively.

Wait, but isn’t this success? Didn’t we make the agent have no trajectory length preference?

No. Suppose:

Same as before, but now there’s a little guy standing by the End/Continue button. Sometimes he likes to press button randomly.

Do we kill the guy? Yes we certainly do, he will mess up our careful plans.