I feel like I want a cost in this model, explicitly or implicitly. One intuition: if “reward seeking” generalizes super well and “local skill” doesn’t then we should converge to reward seeking no matter how much weight we put on it initially, because skill is costly and with reward seeking we can get the same level of performance much more cheaply. Another: if reward seeking is basically parasitic on local skill in each environment it shouldn’t be learned (sort of a model with near 1), because it’s not worth paying the cost for the reward seeking skill instead of just further developing local skill.
Also reward seeking or “local skill” could be cheaper within individual environments, though this may be an extension that isn’t super interesting to the question at hand.
I wondered whether “generalizability across environments” works as a definition of what separates reward seeking from local skill but I don’t think so. For example, arithmetic probably generalizes quite well and is not reward seeking.
I feel like I want a cost in this model, explicitly or implicitly. One intuition: if “reward seeking” generalizes super well and “local skill” doesn’t then we should converge to reward seeking no matter how much weight we put on it initially, because skill is costly and with reward seeking we can get the same level of performance much more cheaply. Another: if reward seeking is basically parasitic on local skill in each environment it shouldn’t be learned (sort of a model with near 1), because it’s not worth paying the cost for the reward seeking skill instead of just further developing local skill.
Also reward seeking or “local skill” could be cheaper within individual environments, though this may be an extension that isn’t super interesting to the question at hand.
I wondered whether “generalizability across environments” works as a definition of what separates reward seeking from local skill but I don’t think so. For example, arithmetic probably generalizes quite well and is not reward seeking.
I agree, I am guessing most of the abstract reasoning skills that we are ideally looking for fit this generality without being hacky.