I agree that in the limit training the AI to act human-like is going to break down, and that a superintelligence can’t be particularly human-like even by definition. However, I’m not sure this matters.
The question is how far the “domain of anthropomorphism” will stretch. I think it’s obvious that it could stretch far enough to get something which would be transformative to the extent that the AIs would be better at almost all intellectual tasks that humans (inc. e.g. long-term-planning), on the basis that there exist actual human people such that if there were millions of copies of them running tirelessly at increased speed and with perfect coordination they could do this; Amodei’s “country of geniuses in a datacenter”. Whether you think it actually will is maybe where we disagree. I don’t think we’ve already gone beyond that domain, because while existing AIs are “superhuman” in some respects they act extremely human-like in general (perhaps this turns on how close you think “close” is to acting human-like). My expectation is that we probably also won’t go beyond it until some time after we have “AGI” (whatever that means); I agree that after that point we’ll have to come up with a new strategy, but I think that “we” in that case will mean the millions of AI instances dedicated to solving alignment issues. Since I’m only interested in strategies that humans will design and implement, I consider that largely out-of-scope.
If it’s true that targeting human-like behaviour is impossible to do while training for long-term thinking ability and it’s not realistic to get it “back on track” before human-level AI, I agree that that would mean that it would be correct to consider more strongly alignment strategies that would otherwise compromise human-like behaviour (if not corrigibility specifically).
The question is how far the “domain of anthropomorphism” will stretch.
The domain of anthropomorphism.
I suspect LLM’s are the mind equivalent of those robots with realistic silicone faces. Humans have a strong tendency to anthropomorphize. We see faces in clouds. The LLM’s are trained in a way that rewards a superfical humanlike appearance.
I agree that in the limit training the AI to act human-like is going to break down, and that a superintelligence can’t be particularly human-like even by definition. However, I’m not sure this matters.
The question is how far the “domain of anthropomorphism” will stretch. I think it’s obvious that it could stretch far enough to get something which would be transformative to the extent that the AIs would be better at almost all intellectual tasks that humans (inc. e.g. long-term-planning), on the basis that there exist actual human people such that if there were millions of copies of them running tirelessly at increased speed and with perfect coordination they could do this; Amodei’s “country of geniuses in a datacenter”. Whether you think it actually will is maybe where we disagree. I don’t think we’ve already gone beyond that domain, because while existing AIs are “superhuman” in some respects they act extremely human-like in general (perhaps this turns on how close you think “close” is to acting human-like). My expectation is that we probably also won’t go beyond it until some time after we have “AGI” (whatever that means); I agree that after that point we’ll have to come up with a new strategy, but I think that “we” in that case will mean the millions of AI instances dedicated to solving alignment issues. Since I’m only interested in strategies that humans will design and implement, I consider that largely out-of-scope.
If it’s true that targeting human-like behaviour is impossible to do while training for long-term thinking ability and it’s not realistic to get it “back on track” before human-level AI, I agree that that would mean that it would be correct to consider more strongly alignment strategies that would otherwise compromise human-like behaviour (if not corrigibility specifically).
The domain of anthropomorphism.
I suspect LLM’s are the mind equivalent of those robots with realistic silicone faces. Humans have a strong tendency to anthropomorphize. We see faces in clouds. The LLM’s are trained in a way that rewards a superfical humanlike appearance.