interstice comments on Introducing Corrigibility (an FAI research subfield)

interstice 24 Oct 2014 3:13 UTC
2 points
How does this differ from indifference?
- Karl 24 Oct 2014 4:00 UTC
  3 points
  Parent
  In the indifference formalism the agent in selecting A1 act like a UN agent that believe that the shutdown button will not be pressed, therefore it create perverse incentives to “manage the news”. Which means that if the agent can cause his shutdown button to be pressed in the event of bad news, it will.
  
  My formulation avoid this pathological behavior by instead making the agent select A1 as if it was a UN-agent which believed that it would continue to optimize according to UN even in the vent of the button being pressed which avoid the perverse incentives to “manage the news”, while still not having any incentives to avoid the button being pressed because the agent will act like it believe that pressing the button will not cause it to initiate a shutdown.