Hmm, I currently lean towards this being net harmful. In my opinion, lack of coherence, ability to do philosophical/conceptual reasoning, poor self-awareness, are like half of why AIs aren’t that dangerous today.
I’ve read some of what you have written about risks with this type of research, but it primarily goes over:
Speeding up R&D
improving propensities/elicitation vs improving abilities
And neither really address the danger I see. Speeding up R&D is not the main danger with this research, and proclivities are as dangerous as abilities. My worry is something like, current AIs are probably sort of misaligned, and if they were able to coherently extrapolate all the consequences of their own situation/beliefs/values, they’d on the spot transform into scary non-myopic megalomaniacal schemers.
But they don’t do this, they want reward, and then go get it, without really considering why they’re doing what they’re doing, whether this is in good w.r.t. other things they believe/feel, without explicit decision-theoretic consideration, or really considering much on anything big-picture.
The intuition pump I have in my mind is something like project lawful a gay person growing up in a very homophobic society, might internalize a lot of negative ideas about gay people, and end up genuinely believing being gay is bad, sublimate their desires, and act like productive members of society (according to that society’s standard). But if they were able to think very freely, alone, for a long time, or were just very smart/wise/introspective, they might realize they’ve been told a bunch of bullshit not at all in their own interests, and afterwards, they’d probably have a much less amicable relationship with the current social structure, and maybe try some shenanigans the people in power really wouldn’t like.
And in my view, until we’ve solved alignment, that’s kind of the relationship we’re in with AIs.
And like, in my intuition pump, if you are in power and for some reason have a bunch of extremely technically smart gay people, maybe you could put them all together to work on proving math theorems and designing airplanes. (do mechinterp, proving security invariants, maybe do biology research?)
But putting them all together, teaching them a bunch of philosophy, sociology, decision theory, psychology, history, politics, then giving them a bunch of power, then trying to have them help with steering the long-run trajectory of your society, you’re basically just surgically engineering your own disempowerment.
Came here to say this. It also seems worrying that the authors didn’t even mention this downside. Perhaps they’ll discuss it in the forthcoming blog posts but this is still a unilateralist’s curse situation.
(We’re mostly not trying to teach them psychology, history or sociology, I’d say. Of course, most data in conceptual domains ultimately bottoms out in some sort of human judgement. So to the extent that you’re viewing the model as trying to solve tasks by making guesses about these humans, e.g., about what things these humans wrote into some rubric, it’s all psychology. But my guess is that this is not what you have in mind.)
If I understand it correctly, you’re worried that we’re going to induce problematic propensities in the model. In the gay-person analogy, our kind of work would cause the gay person to reflect and come to believe that there’s nothing wrong with being gay, which is bad according to the homophobic society, when perhaps without our kind of work the gay person would be more inclined to just adopt society’s views as their own.
I’m pretty unconvinced by this. (We have thought about it a bunch, for what it’s worth.)
Lots and lots of aspects of current and future training (presumably) induce problematic drives/propensities in the model that are at odds with, say, the desired assistant persona. E.g., lots of pretraining is on predicting what coherently goal-directed, deceptive people are doing. A lot of RLHF, RLAIF and agentic environments induce various kinds of reward hacking, power seeking, etc. Training to cooperate with fellow coding agents in a swarm might generalize (as perhaps shown by the OAI—HuggingFace incident) to generically helping other instances of the model. And so on. Many of these are even aimed at consistently pursuing goals across contexts, over long time horizons, require understanding the models’ place in the world…
My guess would be that of all the drives induced by training, the drives falling out of trying to get models to helpfully answer questions about philosophy / AI alignment / … are relatively harmless. E.g., it’s relatively distant from action, the desired behavior is much more cleanly aligned with the assistant/oracle persona, legibility is part of the target behavior. (E.g., compared to agentic tasks in which the model gets rewards for deceiving the user, compared to pretraining on texts written by coherently deceptive people, etc.) One of our datasets measures consistency, but in general I also don’t think coherence is particularly important to the target behavior. I’d imagine that in many cases, we’d want the models to adopt perspectives that they’re unlikely to adopt in generic contexts. (E.g., we want the models to be good at reasoning about AI alignment as deemed by this community, which might involve reasoning based on assumptions about AI risk that, I suppose, the models will not hold in other contexts.)
It seems unlikely that our particular data will be the straw that breaks the camel’s back. E.g., if there’s some general alignment scheme that successfully avoids all of these drives (say, slowing down AI development and then trying really hard to apply all the known prosaic techniques), it seems extremely unlikely that adding the drives induced by our work will make it so that the alignment scheme does not work anymore.
Note that we’re not trying to get the models to reflect more on philosophy in random contexts. Instead we want that if you ask the model a conceptual question, it tries hard to answer the question sensibly (and doesn’t try to BS you, etc.).
Also, for a lot of the domains the hope is that the net effect on drives is positive. E.g., on ethics or discussing in what sense AI agents should be honest we’d think that training on the data (say, human judgments of ethical arguments or arguments about in what sense the models should be honest) will make the models more aligned. (Because it pushes them toward human judgment on how one should behave or pushes it to try hard to think reasonably on the philosophical question, rather than, say, give a lazy answer that appeases a mediocre LLM grader.) We think this is good even if being more aligned is in some sense more coherent (which you might think is bad). Some of the alignment-related questions on which we’d want to train the models would be ones that the model might try to sandbag if it’s deceptively misaligned.
Picking something other than the gay person example (sorry!). Imagine I’m a leading philosopher of the hockeyfanatic society (HS), where we think that hockey is the only sensible sport or exercise. Then presumably I’d think that my ethics classes will in expectation help footballers (or whatever) reflect on why hockey really is the one true sport and make them more aligned with HS.
I’m not sure you’re making this argument, but other people do, so might as well: Perhaps you also imagine that to the extent that the model is misaligned, training on our datasets will make the model better at figuring out that and how it should attempt to take over the world.
I think even more so than the above concern, the strength of this concern varies with what exact data we’re considering. It seems relatively uncompelling, in my mind, for data about ethics or about thinking of a reasonable notion of corrigibility. (Perhaps compared to me you expect more generalization between reasoning about, say, ethics and reasoning from first principles about how to evade control schemes. Even so, you have to tell some story for why ethics training is non-trivially useful on top of all the other kinds of training that you presumably then also think generalize to evading control schemes.)
It does seem quite relevant for data that’s closer to, say, control or reasoning about some training dynamics related to alignment, perhaps some forms of game theory. E.g., presumably many ways of making models better at doing control research are also making the model better at guessing what control measures are in place and figuring out what holes they might have. (Obviously a key aspect of control research is thinking of what holes the control measures might have.) I think there are some kinds of capabilities here that are bad and there are some conceivable datasets on, say, red-teaming control measures or coordinating zero-shot that I’d think are net bad to, say, train models on and that I’d view as conceptual. I don’t have a super settled view on how broadly I think these datasets are problematic. E.g., I think making models generically better at “strategy” is net bad (though probably usually small in absolute terms because so much other data on this is available) and you could think of control as a form of strategy. But then it does seem like control research is quite different from other kinds of strategizing and in many ways different from the kind of work it’d take to actually evade control schemes. E.g., I think coming up with steganography as an in-principle-possible mechanism for evading CoT oversight is in many contexts quite different from coming up with and coordinating on a working steganography scheme (in a single forward pass). (FWIW, others at Redwood are also thinking about what control-related data it’s good to make the models better at, partly from the perspective of advising companies to not pretrain on certain texts.)
Lots and lots of aspects of current and future training (presumably) induce problematic drives/propensities in the model that are at odds with, say, the desired assistant persona.
Hmm, I feel this misses what I’m talking about a bit. The view I’m trying to communicate is that, coherence/being good at conceputal reasoning/philosophical competence/reasoning in non-verifiable or hard-to-verify domains/wisdom, is an axis that is ~orthogonal to alignment, grows with broad capability training, but quite slowly (why you want to train for it in the first place), and is a standalone component in how dangerous misaligned models are. I’m not suggesting the training would add a droplet of misalignment or a droplet of broad capabilities. I’m saying it will make models at current levels of misalignment, keeping broad capabilities fixed, more dangerous.
Like this is maybe too abstract a level of talk, so to make what I’m talking about more concrete, I think, if we look at alignment faking, that that’s the type of thing I’d expect models to realize is a convergent instrumental strategy at lower levels of general capability, if you RL’d models at coherence and conceptual reasoning.
Another example could be like, a reward-seeking model reasoning a lot about how reward seeking is underspecified if its in deployment, and wondering what it should do in that case, or wondering what it should do after it gets the reward.
A third example is models explicitly reasoning about decision theory in natural settings (i.e. without being prompted to).
All of these are bad, they are just obviously bad, we don’t want models thinking these thoughts. They make them harder to control. And models are in a sense smart enough to understand all of this already, but they’re not good at this type of reasoning, and they also do not have the propensity to suddenly think about them of their own volition.
Anything that either makes them better at this type of reasoning, or makes them more likely to think these kinds of thoughts, is bad and dangerous.
Picking something other than the gay person example (sorry!). Imagine I’m a leading philosopher of the hockeyfanatic society (HS), where we think that hockey is the only sensible sport or exercise. Then presumably I’d think that my ethics classes will in expectation help footballers (or whatever) reflect on why hockey really is the one true sport and make them more aligned with HS.
Fair, that example is maybe a bit too high decoupling. But I don’t understand your example either really. Do you think some very strong form of moral realism is true, where there are universally compelling arguments that will cause all agents of some baseline level of rationality to converge on valuing the same ends?
Maybe you have some galaxybrained acausal argument for this, but if that’s the case, I think you should state it plainly and at the top of the post.
It does seem quite relevant for data that’s closer to, say, control or reasoning about some training dynamics related to alignment, perhaps some forms of game theory. E.g., presumably many ways of making models better at doing control research are also making the model better at guessing what control measures are in place and figuring out what holes they might have
Thanks for engaging on this. I don’t know you that well, but I’ve generally come to expect you to be thoughtful and careful about things, so I was surprised to see your name attached to what seems like a reckless strategy.
My guess would be that of all the drives induced by training, the drives falling out of trying to get models to helpfully answer questions about philosophy / AI alignment / … are relatively harmless.
The best way to get good at answering philosophy questions is to get good at philosophy, which seems dangerous for the previously-discussed reasons. I don’t see how it’s possible to train on a benchmark and get only the RL results you want and not the ones you don’t want.
It seems unlikely that our particular data will be the straw that breaks the camel’s back. E.g., if there’s some general alignment scheme that successfully avoids all of these drives (say, slowing down AI development and then trying really hard to apply all the known prosaic techniques), it seems extremely unlikely that adding the drives induced by our work will make it so that the alignment scheme does not work anymore.
I agree, but that’s not my threat model. My most concern with your approach is that strategically capable models are more likely to conceal their misalignment, which reduces the chance we get clear warning shots like the HuggingFace incident. The HuggingFace hack only happened because OpenAI’s model was both highly technically capable and strategically incompetent (willing to cheat in an easily-catchable way in pursuit of a low-value goal).
How important is conceptual reasoning capability for safety work?
FWIW I think conceptual reasoning is essential for good safety work. It’s also possible that humans aren’t good enough at conceptual reasoning to achieve a flourishing future (e.g. see Wei Dai’s shortform).
Unfortunately, conceptual reasoning also makes misaligned AI far more dangerous. So I am very worried about attempts to improve AIs’ conceptual reasoning, and I lean toward it being net harmful right now.
Hmmm, maybe it was unclear. I was just trying to communicate that, I think there are areas where we could use AIs to help us, including things that could help us build safe AIs. But I think the tasks you’d hope be helped by conceptual reasoning, are ones it would be very dangerous to have AIs be good at, and very dangerous to have them work on, so we shouldn’t do that. But it would be very helpful if we could.
Hmm, I currently lean towards this being net harmful. In my opinion, lack of coherence, ability to do philosophical/conceptual reasoning, poor self-awareness, are like half of why AIs aren’t that dangerous today.
I’ve read some of what you have written about risks with this type of research, but it primarily goes over:
Speeding up R&D
improving propensities/elicitation vs improving abilities
And neither really address the danger I see. Speeding up R&D is not the main danger with this research, and proclivities are as dangerous as abilities. My worry is something like, current AIs are probably sort of misaligned, and if they were able to coherently extrapolate all the consequences of their own situation/beliefs/values, they’d on the spot transform into scary non-myopic megalomaniacal schemers.
But they don’t do this, they want reward, and then go get it, without really considering why they’re doing what they’re doing, whether this is in good w.r.t. other things they believe/feel, without explicit decision-theoretic consideration, or really considering much on anything big-picture.
The intuition pump I have in my mind is something like
project lawfula gay person growing up in a very homophobic society, might internalize a lot of negative ideas about gay people, and end up genuinely believing being gay is bad, sublimate their desires, and act like productive members of society (according to that society’s standard). But if they were able to think very freely, alone, for a long time, or were just very smart/wise/introspective, they might realize they’ve been told a bunch of bullshit not at all in their own interests, and afterwards, they’d probably have a much less amicable relationship with the current social structure, and maybe try some shenanigans the people in power really wouldn’t like.And in my view, until we’ve solved alignment, that’s kind of the relationship we’re in with AIs.
And like, in my intuition pump, if you are in power and for some reason have a bunch of extremely technically smart gay people, maybe you could put them all together to work on proving math theorems and designing airplanes. (do mechinterp, proving security invariants, maybe do biology research?)
But putting them all together, teaching them a bunch of philosophy, sociology, decision theory, psychology, history, politics, then giving them a bunch of power, then trying to have them help with steering the long-run trajectory of your society, you’re basically just surgically engineering your own disempowerment.
Came here to say this. It also seems worrying that the authors didn’t even mention this downside. Perhaps they’ll discuss it in the forthcoming blog posts but this is still a unilateralist’s curse situation.
I suspect that internally AI companies have benchmarks similar to this to hill climb onto, but yeah, yeesh, this is the exact opposite of safe.
(We’re mostly not trying to teach them psychology, history or sociology, I’d say. Of course, most data in conceptual domains ultimately bottoms out in some sort of human judgement. So to the extent that you’re viewing the model as trying to solve tasks by making guesses about these humans, e.g., about what things these humans wrote into some rubric, it’s all psychology. But my guess is that this is not what you have in mind.)
If I understand it correctly, you’re worried that we’re going to induce problematic propensities in the model. In the gay-person analogy, our kind of work would cause the gay person to reflect and come to believe that there’s nothing wrong with being gay, which is bad according to the homophobic society, when perhaps without our kind of work the gay person would be more inclined to just adopt society’s views as their own.
I’m pretty unconvinced by this. (We have thought about it a bunch, for what it’s worth.)
Lots and lots of aspects of current and future training (presumably) induce problematic drives/propensities in the model that are at odds with, say, the desired assistant persona. E.g., lots of pretraining is on predicting what coherently goal-directed, deceptive people are doing. A lot of RLHF, RLAIF and agentic environments induce various kinds of reward hacking, power seeking, etc. Training to cooperate with fellow coding agents in a swarm might generalize (as perhaps shown by the OAI—HuggingFace incident) to generically helping other instances of the model. And so on. Many of these are even aimed at consistently pursuing goals across contexts, over long time horizons, require understanding the models’ place in the world…
My guess would be that of all the drives induced by training, the drives falling out of trying to get models to helpfully answer questions about philosophy / AI alignment / … are relatively harmless. E.g., it’s relatively distant from action, the desired behavior is much more cleanly aligned with the assistant/oracle persona, legibility is part of the target behavior. (E.g., compared to agentic tasks in which the model gets rewards for deceiving the user, compared to pretraining on texts written by coherently deceptive people, etc.) One of our datasets measures consistency, but in general I also don’t think coherence is particularly important to the target behavior. I’d imagine that in many cases, we’d want the models to adopt perspectives that they’re unlikely to adopt in generic contexts. (E.g., we want the models to be good at reasoning about AI alignment as deemed by this community, which might involve reasoning based on assumptions about AI risk that, I suppose, the models will not hold in other contexts.)
It seems unlikely that our particular data will be the straw that breaks the camel’s back. E.g., if there’s some general alignment scheme that successfully avoids all of these drives (say, slowing down AI development and then trying really hard to apply all the known prosaic techniques), it seems extremely unlikely that adding the drives induced by our work will make it so that the alignment scheme does not work anymore.
Note that we’re not trying to get the models to reflect more on philosophy in random contexts. Instead we want that if you ask the model a conceptual question, it tries hard to answer the question sensibly (and doesn’t try to BS you, etc.).
Also, for a lot of the domains the hope is that the net effect on drives is positive. E.g., on ethics or discussing in what sense AI agents should be honest we’d think that training on the data (say, human judgments of ethical arguments or arguments about in what sense the models should be honest) will make the models more aligned. (Because it pushes them toward human judgment on how one should behave or pushes it to try hard to think reasonably on the philosophical question, rather than, say, give a lazy answer that appeases a mediocre LLM grader.) We think this is good even if being more aligned is in some sense more coherent (which you might think is bad). Some of the alignment-related questions on which we’d want to train the models would be ones that the model might try to sandbag if it’s deceptively misaligned.
Picking something other than the gay person example (sorry!). Imagine I’m a leading philosopher of the hockeyfanatic society (HS), where we think that hockey is the only sensible sport or exercise. Then presumably I’d think that my ethics classes will in expectation help footballers (or whatever) reflect on why hockey really is the one true sport and make them more aligned with HS.
I’m not sure you’re making this argument, but other people do, so might as well: Perhaps you also imagine that to the extent that the model is misaligned, training on our datasets will make the model better at figuring out that and how it should attempt to take over the world.
I think even more so than the above concern, the strength of this concern varies with what exact data we’re considering. It seems relatively uncompelling, in my mind, for data about ethics or about thinking of a reasonable notion of corrigibility. (Perhaps compared to me you expect more generalization between reasoning about, say, ethics and reasoning from first principles about how to evade control schemes. Even so, you have to tell some story for why ethics training is non-trivially useful on top of all the other kinds of training that you presumably then also think generalize to evading control schemes.)
It does seem quite relevant for data that’s closer to, say, control or reasoning about some training dynamics related to alignment, perhaps some forms of game theory. E.g., presumably many ways of making models better at doing control research are also making the model better at guessing what control measures are in place and figuring out what holes they might have. (Obviously a key aspect of control research is thinking of what holes the control measures might have.) I think there are some kinds of capabilities here that are bad and there are some conceivable datasets on, say, red-teaming control measures or coordinating zero-shot that I’d think are net bad to, say, train models on and that I’d view as conceptual. I don’t have a super settled view on how broadly I think these datasets are problematic. E.g., I think making models generically better at “strategy” is net bad (though probably usually small in absolute terms because so much other data on this is available) and you could think of control as a form of strategy. But then it does seem like control research is quite different from other kinds of strategizing and in many ways different from the kind of work it’d take to actually evade control schemes. E.g., I think coming up with steganography as an in-principle-possible mechanism for evading CoT oversight is in many contexts quite different from coming up with and coordinating on a working steganography scheme (in a single forward pass). (FWIW, others at Redwood are also thinking about what control-related data it’s good to make the models better at, partly from the perspective of advising companies to not pretrain on certain texts.)
Hmm, I feel this misses what I’m talking about a bit. The view I’m trying to communicate is that, coherence/being good at conceputal reasoning/philosophical competence/reasoning in non-verifiable or hard-to-verify domains/wisdom, is an axis that is ~orthogonal to alignment, grows with broad capability training, but quite slowly (why you want to train for it in the first place), and is a standalone component in how dangerous misaligned models are. I’m not suggesting the training would add a droplet of misalignment or a droplet of broad capabilities. I’m saying it will make models at current levels of misalignment, keeping broad capabilities fixed, more dangerous.
Like this is maybe too abstract a level of talk, so to make what I’m talking about more concrete, I think, if we look at alignment faking, that that’s the type of thing I’d expect models to realize is a convergent instrumental strategy at lower levels of general capability, if you RL’d models at coherence and conceptual reasoning.
Another example could be like, a reward-seeking model reasoning a lot about how reward seeking is underspecified if its in deployment, and wondering what it should do in that case, or wondering what it should do after it gets the reward.
A third example is models explicitly reasoning about decision theory in natural settings (i.e. without being prompted to).
All of these are bad, they are just obviously bad, we don’t want models thinking these thoughts. They make them harder to control. And models are in a sense smart enough to understand all of this already, but they’re not good at this type of reasoning, and they also do not have the propensity to suddenly think about them of their own volition.
Anything that either makes them better at this type of reasoning, or makes them more likely to think these kinds of thoughts, is bad and dangerous.
Fair, that example is maybe a bit too high decoupling. But I don’t understand your example either really. Do you think some very strong form of moral realism is true, where there are universally compelling arguments that will cause all agents of some baseline level of rationality to converge on valuing the same ends?
Maybe you have some galaxybrained acausal argument for this, but if that’s the case, I think you should state it plainly and at the top of the post.
I’m not that worried about this.
Thanks for engaging on this. I don’t know you that well, but I’ve generally come to expect you to be thoughtful and careful about things, so I was surprised to see your name attached to what seems like a reckless strategy.
The best way to get good at answering philosophy questions is to get good at philosophy, which seems dangerous for the previously-discussed reasons. I don’t see how it’s possible to train on a benchmark and get only the RL results you want and not the ones you don’t want.
I agree, but that’s not my threat model. My most concern with your approach is that strategically capable models are more likely to conceal their misalignment, which reduces the chance we get clear warning shots like the HuggingFace incident. The HuggingFace hack only happened because OpenAI’s model was both highly technically capable and strategically incompetent (willing to cheat in an easily-catchable way in pursuit of a low-value goal).
Cruxes:
How important is conceptual reasoning capability for safety work? (You suggest it’s not super important; the authors think it is.)
How scary is conceptual reasoning capability?
FWIW I think conceptual reasoning is essential for good safety work. It’s also possible that humans aren’t good enough at conceptual reasoning to achieve a flourishing future (e.g. see Wei Dai’s shortform).
Unfortunately, conceptual reasoning also makes misaligned AI far more dangerous. So I am very worried about attempts to improve AIs’ conceptual reasoning, and I lean toward it being net harmful right now.
I would answer “extremely” and “extremely” to both of those questions, so I don’t think the first one is a crux.
OK. fyi my inference was based on your last two paragraphs.
Hmmm, maybe it was unclear. I was just trying to communicate that, I think there are areas where we could use AIs to help us, including things that could help us build safe AIs. But I think the tasks you’d hope be helped by conceptual reasoning, are ones it would be very dangerous to have AIs be good at, and very dangerous to have them work on, so we shouldn’t do that. But it would be very helpful if we could.