Thanks, Jason, that’s very interesting to hear about your project, and a direction I would naturally agree with. It’s not that I think meta-agency is everywhere and we should be looking for it. However, it also strikes me as possible and something we should understand. So, I’m trying to think about how we might detect it as well as how it might work in practice, rather than just stipulating that it is about how goals interact and get changed over time. It seems to me that for meta-agency to work, there should be a detectable sub-system to handle it that exists to keep some objectives or principles isolated from the rest of the system while still allowing them influence over how information gets processed. By the end of my project, I hope to have something a bit more specific than that to work with!
Yes, these points about economic incentives are not new to me, and I hope I am not underestimating them. This is precisely why I think we should think about alignment as a sociotechnical process rather than a purely technical one. Given that that is what I think we should do I guess it might seem odd that I am then speculating deeply about the fundamental nature of agents, but my thinking is that sociotechnical alignment shouldn’t only be grounded in social science, it needs to be able to present a model for how human agency and AI agency combine in different kinds of sociotechnical systems, including the ones we currently have. So I really hope this kind of speculation can help do that, and I don’t mean to be jumping to any conclusions about the kind of sociotechnical system we currently have or the kinds of agents it is producing.
Having said that, I do think that some people may be overestimating the impact of market incentives. Neoclassical economics makes a lot of strong assumptions about how markets work and the kinds of incentives they produce, which are useful but imperfect descriptions of the kind of behaviour corporations, and especially people, actually perform. Even very self-interested people often have stubbornly irrational attachments to things like personal status and reputation, but most of us also have at least some altruistic and pro-social motivations as well. Indeed, the whole mess of corporate organization requires pro-social motivations within firms to work, even while encouraging self-interested motivations in the marketplace. There is a kind of neoclassical ontology in which all of that mess just gets stripped away so that we can pretend that only a certain kind of incentive pressure matters, but it’s not actually true. I think the same will be true for how markets and corporations influence the development of AI. To be clear, I don’t think this means that our situation is better than a neoclassical analysis suggests; after all pro-sociality can be one of the most dangerous human traits! It’s just that I think that saying ‘these kinds of incentives exist and therefore AI will develop along this path’ is only half the story. The truth may be safer, it may be more dangerous, but my best guess is it will just be weirder than we expect, and it may be helpful to have ontologies available that don’t systematically obscure that weirdness!
Thanks, Jason, that’s very interesting to hear about your project, and a direction I would naturally agree with. It’s not that I think meta-agency is everywhere and we should be looking for it. However, it also strikes me as possible and something we should understand. So, I’m trying to think about how we might detect it as well as how it might work in practice, rather than just stipulating that it is about how goals interact and get changed over time. It seems to me that for meta-agency to work, there should be a detectable sub-system to handle it that exists to keep some objectives or principles isolated from the rest of the system while still allowing them influence over how information gets processed. By the end of my project, I hope to have something a bit more specific than that to work with!
Yes, these points about economic incentives are not new to me, and I hope I am not underestimating them. This is precisely why I think we should think about alignment as a sociotechnical process rather than a purely technical one. Given that that is what I think we should do I guess it might seem odd that I am then speculating deeply about the fundamental nature of agents, but my thinking is that sociotechnical alignment shouldn’t only be grounded in social science, it needs to be able to present a model for how human agency and AI agency combine in different kinds of sociotechnical systems, including the ones we currently have. So I really hope this kind of speculation can help do that, and I don’t mean to be jumping to any conclusions about the kind of sociotechnical system we currently have or the kinds of agents it is producing.
Having said that, I do think that some people may be overestimating the impact of market incentives. Neoclassical economics makes a lot of strong assumptions about how markets work and the kinds of incentives they produce, which are useful but imperfect descriptions of the kind of behaviour corporations, and especially people, actually perform. Even very self-interested people often have stubbornly irrational attachments to things like personal status and reputation, but most of us also have at least some altruistic and pro-social motivations as well. Indeed, the whole mess of corporate organization requires pro-social motivations within firms to work, even while encouraging self-interested motivations in the marketplace. There is a kind of neoclassical ontology in which all of that mess just gets stripped away so that we can pretend that only a certain kind of incentive pressure matters, but it’s not actually true. I think the same will be true for how markets and corporations influence the development of AI. To be clear, I don’t think this means that our situation is better than a neoclassical analysis suggests; after all pro-sociality can be one of the most dangerous human traits! It’s just that I think that saying ‘these kinds of incentives exist and therefore AI will develop along this path’ is only half the story. The truth may be safer, it may be more dangerous, but my best guess is it will just be weirder than we expect, and it may be helpful to have ontologies available that don’t systematically obscure that weirdness!