I notice I’m very confused about why current AI models are so terrible at planning and high level reasoning, so bad at zooming out beyond the exact given task. They seem to lack agency based on their values.
Cognitively, I seriously doubt high level planning is inherently that hard. Probably, models just aren’t being trained to do broad planning, only planning “how to complete single tasks over a long time horizon.”
This means we could be living in a huge AI safety overhang, and all it would take is some tweak to training, and suddenly we’d have models forming complex plans and taking independent actions to achieve them. Which is the sort of thing that a lot of old-school alignment talk assumed would happen and showed would be very bad.
I notice I’m very confused about why current AI models are so terrible at planning and high level reasoning, so bad at zooming out beyond the exact given task. They seem to lack agency based on their values.
Cognitively, I seriously doubt high level planning is inherently that hard. Probably, models just aren’t being trained to do broad planning, only planning “how to complete single tasks over a long time horizon.”
This means we could be living in a huge AI safety overhang, and all it would take is some tweak to training, and suddenly we’d have models forming complex plans and taking independent actions to achieve them. Which is the sort of thing that a lot of old-school alignment talk assumed would happen and showed would be very bad.