Thanks for the response, this all makes a lot of sense and is very interesting.
Edward and I’s agenda is taking a fairly agnostic stance (at least to begin with) on what sorts of internal structure AI cognition has, and mostly aims at exploring and evaluating what kinds of structures seem likely or are useful to reason via. Modelling meta-agency a certain way or modelling agents as having agentic shards might help our dev-cog models fit the data better (including better predictions for out-of-distribution training pipelines), then this provides good evidence that these are useful structures for reasoning about agency, but we are not committed to them. In general we are just trying to seek out the most useful structures, whatever they may be. On that note, we did have some pre-preliminary evidence a while back for some more shard-like cognitive models, but for now we’ve found very simple structures (with no modelled meta-agency) to be good enough. I imagine as we scale the agenda in what we try to predict and what dev-cog models work well, if a meta-agentic-shaped gap starts appearing we can then start to properly tackle it.
On your points on Vanessa Kosoy and Veedracs, I broadly agree, but I think perhaps you are overestimating the degree of possible co-ordination on AI, and underestimating the effects of incentive pressures, relative to other people working on AI Safety. If I’m wrong about this and the following arguments / points are not new to you I apologise. Essentially, there are economic pressures to build strong goal-oriented AI. Powerful AI would be able to create lots of wealth autonomously running companies, or trading stocks (and actually now has). Because of this, AI researchers or the AI safety community do not get to dictate what the nature of AI should be (without sufficient co-ordination that seems very unlikely). Instead, we have to try and design a safe approach for AI that will itself outcompete or prevent the sorts of AIs that might be built due to humans just following the incentives to build CEO-bot. I think this is mostly what Vanessa is gesturing at when she says we need to be worried about systems with bad goals. It’s not “well some AI safety people might make AIs with bad goals so we need to be robust to that”, in which the argument you’re making of “hey maybe this is a bad framing for AI safety people to make” makes sense, but instead it’s “some non-safety-minded people might make CEO-bot and so we need to be robust to that / win in worlds where that happens”. CEO-bot is one specific example, there are other forms this AI might take, but I think the general point of there existing incentive pressures to create goal-oriented AI systems by people who might not care or understand about AI safety is enough to motivate AI safety people not being able to completely ignore the space of AI minds that are somewhat goal-oriented. I think Evan Hubinger explains / illustrates this point really well in his somewhat old post on 11 proposals for safe AI, where he uses training and performance competitiveness as two of his four desiderata. Hope that helps!
Thanks, Jason, that’s very interesting to hear about your project, and a direction I would naturally agree with. It’s not that I think meta-agency is everywhere and we should be looking for it. However, it also strikes me as possible and something we should understand. So, I’m trying to think about how we might detect it as well as how it might work in practice, rather than just stipulating that it is about how goals interact and get changed over time. It seems to me that for meta-agency to work, there should be a detectable sub-system to handle it that exists to keep some objectives or principles isolated from the rest of the system while still allowing them influence over how information gets processed. By the end of my project, I hope to have something a bit more specific than that to work with!
Yes, these points about economic incentives are not new to me, and I hope I am not underestimating them. This is precisely why I think we should think about alignment as a sociotechnical process rather than a purely technical one. Given that that is what I think we should do I guess it might seem odd that I am then speculating deeply about the fundamental nature of agents, but my thinking is that sociotechnical alignment shouldn’t only be grounded in social science, it needs to be able to present a model for how human agency and AI agency combine in different kinds of sociotechnical systems, including the ones we currently have. So I really hope this kind of speculation can help do that, and I don’t mean to be jumping to any conclusions about the kind of sociotechnical system we currently have or the kinds of agents it is producing.
Having said that, I do think that some people may be overestimating the impact of market incentives. Neoclassical economics makes a lot of strong assumptions about how markets work and the kinds of incentives they produce, which are useful but imperfect descriptions of the kind of behaviour corporations, and especially people, actually perform. Even very self-interested people often have stubbornly irrational attachments to things like personal status and reputation, but most of us also have at least some altruistic and pro-social motivations as well. Indeed, the whole mess of corporate organization requires pro-social motivations within firms to work, even while encouraging self-interested motivations in the marketplace. There is a kind of neoclassical ontology in which all of that mess just gets stripped away so that we can pretend that only a certain kind of incentive pressure matters, but it’s not actually true. I think the same will be true for how markets and corporations influence the development of AI. To be clear, I don’t think this means that our situation is better than a neoclassical analysis suggests; after all pro-sociality can be one of the most dangerous human traits! It’s just that I think that saying ‘these kinds of incentives exist and therefore AI will develop along this path’ is only half the story. The truth may be safer, it may be more dangerous, but my best guess is it will just be weirder than we expect, and it may be helpful to have ontologies available that don’t systematically obscure that weirdness!
Thanks for the response, this all makes a lot of sense and is very interesting.
Edward and I’s agenda is taking a fairly agnostic stance (at least to begin with) on what sorts of internal structure AI cognition has, and mostly aims at exploring and evaluating what kinds of structures seem likely or are useful to reason via. Modelling meta-agency a certain way or modelling agents as having agentic shards might help our dev-cog models fit the data better (including better predictions for out-of-distribution training pipelines), then this provides good evidence that these are useful structures for reasoning about agency, but we are not committed to them. In general we are just trying to seek out the most useful structures, whatever they may be. On that note, we did have some pre-preliminary evidence a while back for some more shard-like cognitive models, but for now we’ve found very simple structures (with no modelled meta-agency) to be good enough. I imagine as we scale the agenda in what we try to predict and what dev-cog models work well, if a meta-agentic-shaped gap starts appearing we can then start to properly tackle it.
On your points on Vanessa Kosoy and Veedracs, I broadly agree, but I think perhaps you are overestimating the degree of possible co-ordination on AI, and underestimating the effects of incentive pressures, relative to other people working on AI Safety. If I’m wrong about this and the following arguments / points are not new to you I apologise. Essentially, there are economic pressures to build strong goal-oriented AI. Powerful AI would be able to create lots of wealth autonomously running companies, or trading stocks (and actually now has). Because of this, AI researchers or the AI safety community do not get to dictate what the nature of AI should be (without sufficient co-ordination that seems very unlikely). Instead, we have to try and design a safe approach for AI that will itself outcompete or prevent the sorts of AIs that might be built due to humans just following the incentives to build CEO-bot. I think this is mostly what Vanessa is gesturing at when she says we need to be worried about systems with bad goals. It’s not “well some AI safety people might make AIs with bad goals so we need to be robust to that”, in which the argument you’re making of “hey maybe this is a bad framing for AI safety people to make” makes sense, but instead it’s “some non-safety-minded people might make CEO-bot and so we need to be robust to that / win in worlds where that happens”. CEO-bot is one specific example, there are other forms this AI might take, but I think the general point of there existing incentive pressures to create goal-oriented AI systems by people who might not care or understand about AI safety is enough to motivate AI safety people not being able to completely ignore the space of AI minds that are somewhat goal-oriented. I think Evan Hubinger explains / illustrates this point really well in his somewhat old post on 11 proposals for safe AI, where he uses training and performance competitiveness as two of his four desiderata. Hope that helps!
Thanks, Jason, that’s very interesting to hear about your project, and a direction I would naturally agree with. It’s not that I think meta-agency is everywhere and we should be looking for it. However, it also strikes me as possible and something we should understand. So, I’m trying to think about how we might detect it as well as how it might work in practice, rather than just stipulating that it is about how goals interact and get changed over time. It seems to me that for meta-agency to work, there should be a detectable sub-system to handle it that exists to keep some objectives or principles isolated from the rest of the system while still allowing them influence over how information gets processed. By the end of my project, I hope to have something a bit more specific than that to work with!
Yes, these points about economic incentives are not new to me, and I hope I am not underestimating them. This is precisely why I think we should think about alignment as a sociotechnical process rather than a purely technical one. Given that that is what I think we should do I guess it might seem odd that I am then speculating deeply about the fundamental nature of agents, but my thinking is that sociotechnical alignment shouldn’t only be grounded in social science, it needs to be able to present a model for how human agency and AI agency combine in different kinds of sociotechnical systems, including the ones we currently have. So I really hope this kind of speculation can help do that, and I don’t mean to be jumping to any conclusions about the kind of sociotechnical system we currently have or the kinds of agents it is producing.
Having said that, I do think that some people may be overestimating the impact of market incentives. Neoclassical economics makes a lot of strong assumptions about how markets work and the kinds of incentives they produce, which are useful but imperfect descriptions of the kind of behaviour corporations, and especially people, actually perform. Even very self-interested people often have stubbornly irrational attachments to things like personal status and reputation, but most of us also have at least some altruistic and pro-social motivations as well. Indeed, the whole mess of corporate organization requires pro-social motivations within firms to work, even while encouraging self-interested motivations in the marketplace. There is a kind of neoclassical ontology in which all of that mess just gets stripped away so that we can pretend that only a certain kind of incentive pressure matters, but it’s not actually true. I think the same will be true for how markets and corporations influence the development of AI. To be clear, I don’t think this means that our situation is better than a neoclassical analysis suggests; after all pro-sociality can be one of the most dangerous human traits! It’s just that I think that saying ‘these kinds of incentives exist and therefore AI will develop along this path’ is only half the story. The truth may be safer, it may be more dangerous, but my best guess is it will just be weirder than we expect, and it may be helpful to have ontologies available that don’t systematically obscure that weirdness!