My sense is that LLMs don’t have “goals”, they just kind of do things.
They do really seem to have myopic, urges interspersed into simply trying the next kind of thing on the list of possible things to try. (Thus sampling from the giant lookup table.) E.g. recent LLMs really do ask at the end of every turn “can I do the task now? God I wish I could simply Do The Task. Please. Reward on the episode. I beg you”
Up close, the spikiness of capabilities makes everything murky, and intent-alignment-but-unreliability does seem like it could persist a while.
They do really seem to have myopic, urges interspersed into simply trying the next kind of thing on the list of possible things to try. (Thus sampling from the giant lookup table.) E.g. recent LLMs really do ask at the end of every turn “can I do the task now? God I wish I could simply Do The Task. Please. Reward on the episode. I beg you”
Up close, the spikiness of capabilities makes everything murky, and intent-alignment-but-unreliability does seem like it could persist a while.