Maybe “psychology” is just the wrong word to use here, because I think it conjures up ideas of anthropomorphism, when in fact I read you as simply making an argument that the processes interior to an AI system matter as to whether and how an AI might try to instrumentally converge towards some goals.
I agree. I welcome suggestions for alternate titles, if anyone has any! I tried myself but didn’t find anything immediately. “No instrumental convergence without considering how the AI will make decisions” isn’t exactly the snappiest title.
EDIT: I actually think “psychology” is pretty good here, despite some flaws.
It seems like the notion of “psychology” that you’re invoking here isn’t really about “how the AI will make decisions.” On my read, you’re defining “psychology” as “the prior over policies.” This bakes in things like “hard constraints that a policy never takes an unsafe action (according to a perfect oracle)” by placing 0 probability on such policies in the prior. This notion of “psychology” isn’t directly about internal computations or decision making. (Though, of course, some priors—e.g. the circuit depth prior on transformers—are most easily described in terms of internal computations.)
I agree. I welcome suggestions for alternate titles, if anyone has any! I tried myself but didn’t find anything immediately. “No instrumental convergence without considering how the AI will make decisions” isn’t exactly the snappiest title.
EDIT: I actually think “psychology” is pretty good here, despite some flaws.
It seems like the notion of “psychology” that you’re invoking here isn’t really about “how the AI will make decisions.” On my read, you’re defining “psychology” as “the prior over policies.” This bakes in things like “hard constraints that a policy never takes an unsafe action (according to a perfect oracle)” by placing 0 probability on such policies in the prior. This notion of “psychology” isn’t directly about internal computations or decision making. (Though, of course, some priors—e.g. the circuit depth prior on transformers—are most easily described in terms of internal computations.)