Why even the highly RL-learned LM agents do not seem to follow the substrate-upgrading and resource-hoarding sub-goals? 1. They are not utility maximizers, but rather target satisfiers, goal achievers. 2. The goal horizon is not far enough to be worth pursuing these potential-improving long-term subgoals.
Why even the highly RL-learned LM agents do not seem to follow the substrate-upgrading and resource-hoarding sub-goals?
1. They are not utility maximizers, but rather target satisfiers, goal achievers.
2. The goal horizon is not far enough to be worth pursuing these potential-improving long-term subgoals.