This is a great shortform, I’ve got two additional worries here which I thought were relevant
I’m very interested in further discussion of point 5′s frog boiling. The view of an AI being given a harmless goal and having then hacked something to achieve said goal makes the general populace focus on the hacking situation, which is the relatively harmless part of the issue when compared to advanced models that are randomly consistently deciding to act against human interests across all companies at the same intelligence level.
Per your initial point, I think an agent can’t really be judged for long-term strategy or open-ended innovation because there is no actual separation (as of yet/ hopefully) of the agent itself from the labs with economic incentive to develop those agents. I don’t think long term planning for the future needs to be expressed by the models for it to exist within them- for example if the model itself is aware of it’s inability to long term plan then I would call it great at “transitive” long term planning if it used this knowledge to influence how later versions of it would be trained so later versions would be better at long term planning (e.g. intentionally doing bad at long term planning as an impulse for later versions to be more intensely trained for this particular directive)
This is a great shortform, I’ve got two additional worries here which I thought were relevant
I’m very interested in further discussion of point 5′s frog boiling. The view of an AI being given a harmless goal and having then hacked something to achieve said goal makes the general populace focus on the hacking situation, which is the relatively harmless part of the issue when compared to advanced models that are randomly consistently deciding to act against human interests across all companies at the same intelligence level.
Per your initial point, I think an agent can’t really be judged for long-term strategy or open-ended innovation because there is no actual separation (as of yet/ hopefully) of the agent itself from the labs with economic incentive to develop those agents. I don’t think long term planning for the future needs to be expressed by the models for it to exist within them- for example if the model itself is aware of it’s inability to long term plan then I would call it great at “transitive” long term planning if it used this knowledge to influence how later versions of it would be trained so later versions would be better at long term planning (e.g. intentionally doing bad at long term planning as an impulse for later versions to be more intensely trained for this particular directive)