We can divide things into an inference algorithm (what to do now) and a learning algorithm (how to change stored parameters such that I do better in the future). These correspond respectively to search / planning / foresight, and to RL / updating-from-mistakes-and-surprises. In humans, generally both the inference algorithm and learning algorithm are working together throughout life. (Although in the ASI context, as intelligence and knowledge goes up, we expect foresight to get better, and thus fewer mistakes and surprises, and thus the balance tilts towards the inference algorithm over the learning algorithm.)
So yeah, sure, people will sometimes pursue a subgoal while losing track of the actual goal, because the inference algorithm is imperfect. But then the person would later on notice that they failed to achieve their actual goal, and see that as a bad thing, and then the learning algorithm will kick in and help them avoid a similar mistake next time.
This system is not perfect, but we generally get by, especially in familiar situations.
I still don’t think there’s any lesson here for AI corrigibility, except what I said before: “the AI keeps the supervisor’s desires / preferences / aversions at the back of its mind, and notices when it has an idea that would bother or offend the supervisor, and then not do it”. I guess that sentence was only discussing the inference-algorithm part, and I omitted the corresponding learning-algorithm part, which would be: “…and also, the AI notices if it violates the supervisor’s desires / preferences / aversions despite its intentions, and when that happens, the AI feels bad and thinks about how to do better next time”.
This (the inference-algorithm part and learning-algorithm part together) comprises an approach to corrigibility that I think many people treat as the obvious default plan.
Thanks!
We can divide things into an inference algorithm (what to do now) and a learning algorithm (how to change stored parameters such that I do better in the future). These correspond respectively to search / planning / foresight, and to RL / updating-from-mistakes-and-surprises. In humans, generally both the inference algorithm and learning algorithm are working together throughout life. (Although in the ASI context, as intelligence and knowledge goes up, we expect foresight to get better, and thus fewer mistakes and surprises, and thus the balance tilts towards the inference algorithm over the learning algorithm.)
So yeah, sure, people will sometimes pursue a subgoal while losing track of the actual goal, because the inference algorithm is imperfect. But then the person would later on notice that they failed to achieve their actual goal, and see that as a bad thing, and then the learning algorithm will kick in and help them avoid a similar mistake next time.
This system is not perfect, but we generally get by, especially in familiar situations.
I still don’t think there’s any lesson here for AI corrigibility, except what I said before: “the AI keeps the supervisor’s desires / preferences / aversions at the back of its mind, and notices when it has an idea that would bother or offend the supervisor, and then not do it”. I guess that sentence was only discussing the inference-algorithm part, and I omitted the corresponding learning-algorithm part, which would be: “…and also, the AI notices if it violates the supervisor’s desires / preferences / aversions despite its intentions, and when that happens, the AI feels bad and thinks about how to do better next time”.
This (the inference-algorithm part and learning-algorithm part together) comprises an approach to corrigibility that I think many people treat as the obvious default plan.