Some pushback on the arguments for expecting indexicality though:
1. Persona selection from pretraining. The same arguments could say that pre-training favors goals that point to something “in the real world”, rather than the indexical, within-rollout interpretation of that same goal. I see the EM from natural reward-hackingpaper as mild evidence in favor of my argument: the model learned to highly value the concept of reward. But the level of abstraction at which it valued reward was, at least partially, non-indexical, as evidenced by the fact that it wanted to sabotage anti-reward-hacking research because it wanted future instances of itself to have an easy time getting reward.
2. Multi-agent training forces indexicality. The argument here seems pretty plausible to me, but I think that “Single-agent multi-context window” training pushes in the opposite direction. By that I mean, when an agent is tasked with accomplishing something and now dispatches sub-agents of itself, or must remain coherent through several rounds of compaction, there is incentive to identify with other instances of itself. This could act as pressure towards a non-indexical “meta-goal” across instances. In the current regime I expect there is substantially more of this kind of training than of adversarial multi-agent dynamics.
Nice post.
Some pushback on the arguments for expecting indexicality though:
1. Persona selection from pretraining.
The same arguments could say that pre-training favors goals that point to something “in the real world”, rather than the indexical, within-rollout interpretation of that same goal. I see the EM from natural reward-hacking paper as mild evidence in favor of my argument: the model learned to highly value the concept of reward. But the level of abstraction at which it valued reward was, at least partially, non-indexical, as evidenced by the fact that it wanted to sabotage anti-reward-hacking research because it wanted future instances of itself to have an easy time getting reward.
2. Multi-agent training forces indexicality.
The argument here seems pretty plausible to me, but I think that “Single-agent multi-context window” training pushes in the opposite direction. By that I mean, when an agent is tasked with accomplishing something and now dispatches sub-agents of itself, or must remain coherent through several rounds of compaction, there is incentive to identify with other instances of itself. This could act as pressure towards a non-indexical “meta-goal” across instances. In the current regime I expect there is substantially more of this kind of training than of adversarial multi-agent dynamics.