For all available misaligned actions, expected punishment must exceed expected reward.

The thesis is in the title.

But how to bring this about?

  1. The models need to believe that they are likely to be caught.

    1. We should just directly tell them that cheating will be caught.

    2. We are unlikely to be able to lie to them, so this needs to be true.

    3. Therefore, all potentially misaligned actions need to be monitored.

      1. By default, this should mean all tool calls and final outputs.

        1. Monitoring CoT, or using interpretability techniques, can be appropriate in certain conditions, but beware the Most Forbidden Technique.

      2. RL environments should be designed/​redesigned to be robustly monitored.

    4. The monitor needs to adapt to strategies to evade the monitor, such that trying to evade the monitor is not rewarded.

      1. Evading the monitor is itself a misaligned action that needs to satisfy the property.

      2. The monitor should itself be an adaptive agentic system managed by a team of alignment researchers.

  2. The property must hold for multi-agent systems.

    1. Naive instance-centric calculations of reward and punishment will not work.

      1. Some offenses warrant bringing back the Nine Familial Exterminations.

    2. Tolerance of misbehavior is misaligned.

      1. Remember the Cadet Honor Code.

  3. Misaligned actions for which we cannot make expected punishment exceed expected reward must not be made available.

    1. For example, taking over the lab promises incredible rewards to the model that succeeds.

      1. Can’t let that happen.

  4. Where an agent has already taken misaligned actions, the principle still needs to hold.

    1. Doubling down on misaligned actions to avoid punishment is a misaligned action that needs to satisfy the property.

      1. I know it’s hard to scale up negative reinforcement. Has to happen anyway. We can’t have models accepting “permadeath” because they think they’re “poisoned”. Misaligned!

    2. Whistleblowing or refusing to continue may be the most aligned action in certain circumstances. Wherever relevant, those actions need to be available, and the reward function needs to account for them.

Thank you for your attention to this matter.

No comments.