There is an idea floating around in the rough shape of “we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better.” The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning.
I feel nervous about this for two reasons. The first is that it’s plausible that AI safety and AI R&D are bottlenecked by many of the same factors: AIs have poor epistemics, are bad at messy conceptual reasoning, and are unreliable at tasks without ground truth. Speeding up progress in any of these areas seems likely to speed up general AI R&D, giving everyone else less time to execute time-bottlenecked agendas (e.g., trying to do Plan A).
The second reason I don’t feel good about this is because I’m less confident it will help make handoff/deference/superalignment go well. To hand off conceptual alignment research to AIs we need to trust them to 1. be good at this research and 2. be generally trustworthy/aligned. We still don’t know how to reliably prevent prosaic outer misalignment issues (e.g., sycophancy or going off-constitution), let alone worse issues that will make AIs less trustworthy in the coming years. For this reason, I expect #2 to be more of a bottleneck to high-stakes alignment research than #1, which seems to be more likely to be an emergent property of more capable models.
I haven’t seen anyone clearly write up these arguments and I think that more public dialog on dual-use AI safety work would be beneficial. In this post I:
Give examples of arguments people have made for the acceleration of alignment-relevant capabilities
Give arguments for why accelerating capabilities differentially useful for doing alignment research shouldn’t be a current priority
Address two counterarguments:
“But what about things like philosophy?”
“But isn’t it pretty unlikely that you help labs make real capabilities progress?”
Examples of arguments for the acceleration of alignment-relevant capabilities
I’m including these examples to give a sense of the kinds of claims being made and what motivates them, not to be comprehensive or fully explain each argument.
Important disclaimer: Note that the rest of the post should not be read as responding to any of these claims specifically (many/most of my arguments don’t apply to all of them) but instead as responding to the general family of arguments shaped like this.
More precisely, our goal is to bring forward in time the point when the capability profile allows for fully automating safety work relative to the point where various even more dangerous capability milestones are reached. (And, broadly speaking, we’d like to avoid accelerating the time at which dangerous capability milestones are reached, though some general acceleration might be unavoidable.)
Although I personally think this is a bad idea, I find it laudable that Greenblatt and Stastny point out the potential tradeoff explicitly. At a high level, they argue that for AIs to be good at alignment research, they need to be broadly aligned, good at AI R&D, and good at messy conceptual things. They also argue that handing off safety research is safer in less capable models, hence the suggestion to differentially accelerate the skills needed for alignment research without pushing general capabilities too much.[1]
If AIs became strategically competent enough, they may realize that RSI is too dangerous because they’re not good enough at alignment or philosophy or strategy, and potentially convince, help, or work with humans to implement an AI pause. This presents an alternative “victory condition” that someone could pursue (e.g. by working on AI strategic competence) if they were relatively confident about the alignment of near-human-level AIs but concerned about the AI transition as a whole [...]
Dai’s argument explicitly states that alignment is a prerequisite for this to work which I agree with. He also explicitly flags that improving strategic competence could make things worse:
But note that if the near-human-level AIs are not aligned, then this effort could backfire by letting them apply better strategy to take over more easily.
Although Dai wants AIs that are good at arranging a pause rather than solving alignment, I argue that it’s analogous to the Greenblatt et Stastny approach. The idea is still that there exists some capability that could be differentially good for safety and we should consider accelerating it.
We’ve honed in on a particular project: Measuring and improving the conceptual reasoning abilities of LLMs through elicitation. Conceptual reasoning is, roughly, reasoning in domains and about questions where we don’t have (access to) ground truth. The prototypical example of this kind of domain is philosophy.
[...] We also believe conceptual research is differentially useful for many other AI safety applications compared to capabilities research and expect the project to have many positive externalities outside of the acausal agenda.
As I understand it, the idea is by making AIs better at conceptual reasoning they will be able to help better at acausal stuff and maybe a wide variety of things that could prevent human extinction or other catastrophe also.
Paul Christiano expresses that advances in agent capabilities could be positive in older writing here but does not say that people should work on this.
Why we should not do this kind of differential acceleration
Alignment bottlenecks are also capabilities bottlenecks
There are certain problems with current AI systems which make them bad at AI R&D:
There are additional problems that make them bad at alignment research:
They can’t reason in domains without ground truth
They are bad at philosophy
Rereading these two lists, they appear to rhyme with each other. To be fair, alignment research does, at a glance, look like the kind of thing that would require more conceptual / philosophical breakthroughs compared to capabilities research. So there are good reasons to think there exist some philosophical skills which could be differentially useful for alignment.
But accelerating reasoning without ground truth broadly seems too bottleneck-y for both to be a good safety target. For other capabilities that appear more narrow, there is the risk of unfavorable generalization. For example, it’s likely that being good at messy reasoning for a somewhat narrow skill requires eliciting more general conceptual skills that apply to other domains. We have seen some examples of this empirically: training models on logic problems can induce new reasoning abilities and generalize to unrelated math benchmarks. Likewise, the “let’s think step by step” technique generalized broadly and motivated new training techniques even though, on the surface, it is simply an elicitation technique for deeper reflection.
At the very least, there is overlap between alignment-capabilities and capabilities-capabilities and, in expectation, some speedup in AI progress conditional on successful differential acceleration.[2] I don’t think it will be trivial to minimize this risk because various stages of the research can be infohazard-rich and even vague details about techniques can be enough to reproduce them.
This could hurt time-bottlenecked strategies
Old-school EAs used to talk about being a force-multiplier. For example, instead of directly distributing bed nets or something you could trade stocks, make millions, then pay the salaries of 10 people who distribute bed nets or work for the bed nets charity. If you simply worked at the charity, you would be only responsible for 1/10th of the bed net impact so you have effectively multiplied yourself.
In AI safety, anything that accelerates AI progress and brings us closer to RSI is effectively a force-minimizer. Shortening timelines gives a large number of projects less time to figure stuff out and lowers the probability that any of them are successful, effectively subtracting workers from those efforts. There are some agendas where the progress on safety-relevant capabilities would speed things up but there are some activities which are bottlenecked by time. Concretely, making timelines shorter means there is less time to advocate for a pause, fundraise for a promising new alignment technique, develop compute verification technology, etc.
Importantly, I would not expect any capabilities acceleration to be differentially useful for things like advocating for a pause. Imagine a world where we have AIs that have the skills that would make them good at lobbying for a pause: because labs will have access to the most compute and the best capabilities (they may not make the best models public), they will have the advantage to advocate for the position they want (which would not be a pause probably).
I don’t think that this would unblock superalignment/hand-off plans
(This is all written under the assumption that superalignment-flavored strategies are a good idea which they may not be. A lot of the time when people say “differentially good for safety” they mean “differentially good for superalignment/handoff” but here I am assuming that this is fine.)
If a future AI system were capable of good alignment research would we trust it to do a good job? I would argue probably not: current AIs can already be sycophantic, dishonest, pretty misaligned, and can go off-persona or off-constitution in undesirable ways. We can expect more outer misalignment failures as we scale up increasingly insane post-training schemes and even possibly inner misalignment problems at sufficient levels of capabilities. In other words, it’s not clear we are on track to have a model that we can trust to hand off our safety research to.
Concretely: working on alignment directly seems just as differentially good for superalignment/handoff compared to any kind of capabilities acceleration evenif we are ignoring the capabilities externalities entirely. (Of course, alignment work can be dual use as well,[3] and for that reason I think it is reasonable to be concerned about alignment research that doesn’t aim to solve the core alignment problem but rather makes incremental progress on product alignment.[4]) The argument for alignment becomes stronger once we account for the externalities and other factors:
Alignment progress may be useful for a wide variety other alignment agendas.
Alignment work has limited capabilities externalities compared to trying to accelerate certain safety-relevant capabilities.
The capabilities bottleneck will likely be solved by the time it matters most: either by default (e.g, with scale you start to get this) or by something more intentional (e.g., downstream of labs trying to train in research taste). Alignment by default is also possible, but conditional on it being true, its unclear if any of this matters because we may just be in a good world.
But what about things like philosophy?
I grant that more narrow capabilities like making models good at philosophy are less likely to be useful for automating AI R&D and probably pretty useful for alignment. But is it possible to develop techniques that only push philosophy and nothing else? If you train on messy philosophical thinking or you get good at eliciting it, would you also be good at more general messy conceptual thinking? Is it even possible to be good at philosophy without being generally really good at messy conceptual thinking? Likewise, if the elicitation technique truly only targets philosophy, could other actors not use a similar technique to elicit more research-relevant conceptual thinking?
Overall, I’m not convinced that narrow capabilities can just be selectively advanced through elicitation or other techniques (or that the same technique could not be applied to additional narrow capabilities).
But isn’t it pretty unlikely that you help labs make real capabilities progress?
(It’s somewhat rare I hear someone make this argument but I have heard it a few times so I thought I would address it.)
I find people can be pretty confident that accelerating safety relevant capabilities like messy conceptual reasoning is really quite tractable and labs are ignoring it. But then when talking about the capabilities externalities or how this could speed up timelines, there is a feeling that pushing the frontier of AI agents is really quite difficult because there is already billions and soon-to-be trillions invested in this and all the low hanging fruit has been picked.
I’m open to the idea that trying to advance capabilities relevant for alignment is intractable and the result of efforts here will be unsuccessful and therefore net-neutral (ignoring the opportunity cost of the money/time spent on the research). But I think it’s hard to hold 1. the research is very tractable, 2. there is overlap with capabilities and 3. any low hanging fruit with capabilities research has already been found. I suppose you could reject #2 but as discussed above, I don’t think we can cleanly split “safety-capabilities” and “capabilities-capabilities.”
I will also note that the open-source frontier may be easier to advance and pretty dangerous also.
My epistemic status
Thinking about superalignment or what actions in the future make it go good/bad requires a certain amount of playing 4D chess with a blindfold on. I think people should be uncertain. I myself am uncertain that superalignment-flavored things are a good idea to begin with. I also don’t think I should be entirely confident that, for example, alignment of AIs capable of alignment research will be a bottleneck to handoff because who knows. The world will be weird.
When I say “differential acceleration of alignment-relevant capabilities is a bad bet” I don’t mean I’m confident it’s negative EV. I’m more saying that this particular bet looks like it takes on a lot of unnecessary risk when there are other things that look equally promising. If the goal is to hand off to AI early-ish (which I’m not claiming is good or bad) then even some prosaic alignment things like trying to systematically understand midtraining or trying to “solve” eval awareness seem less risky while being equally productive and maybe more tractable.
I will say that I feel pretty confident that the dual-use aspects of this kind of research are a real risk that deserves to be more widely discussed and publicly debated. At the very least, it seems good for those pursuing dual-use research to subject themselves to lots of red-teaming (if the idea itself isn’t an infohazard, this should be public) and have a preregistered policy for handling potentially dangerous information.
Greenblatt and Stastny also have ways to address some of the points in my post, but I don’t respond to these specific points directly in an effort to keep the scope broad. (I don’t want to respond to a specific post but instead give reasons for why I don’t find the family of arguments convincing.)
I would argue that even elicitation techniques can be silently risky. For example, “let’s think step by step” seems like somewhat benign elicitation technique that could make the model better at messy philosophical questions. But it also led to reasoning models which was a breakthrough.
I worry that other elicitation techniques could be more broadly applicable than people think because there may be a more general mechanism behind a seemingly narrow capabilities breakthrough.
This is a crux though. If people gave successful examples of narrow capabilities being accelerated in language models where there is no more general mechanism that could be used for other domains, I would update my beliefs. I tend to think that if there is an elicitation technique that is good at surfacing safety-relevant capabilities, there is a good chance it could be repurposed for surfacing arbitrary capabilities.
Working on this misalignment bottleneck can mean a lot of different things and some of those things also have capabilities externalities so we can run into similar problems. RLHF is a standard example of “try to solve outer alignment and accidentally improve general capabilities.”
Things like trying to make the model less eval aware during safety-relevant tests seems robustly good with limited negative externalities. Things like trying to make AIs less reward hacky also falls under the umbrella of “solve outer alignment in early transformative AIs” but I can imagine this making general AI research agents easier to use for non-safety things.
This is all to say that dual-use risks should remain a consideration even when working on things that appear more alignment-flavored.
It’s best to try to solve the core problem but admittedly most alignment research does something closer to “incremental progress to make the alignment metric go up for some model organism on some eval.” I actually generally cautiously support the latter in addition to the former. If we don’t “solve” alignment then we are in a world where we are putting lots of different bandaids on the problem and hoping for the best. This is extremely reckless and people working on bandaid development should understand this and communicate this honestly, but if we are in this world I think we will be pretty grateful that there are different bandaids to choose from.
Differential acceleration of alignment-relevant capabilities is a bad bet
There is an idea floating around in the rough shape of “we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better.” The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning.
I feel nervous about this for two reasons. The first is that it’s plausible that AI safety and AI R&D are bottlenecked by many of the same factors: AIs have poor epistemics, are bad at messy conceptual reasoning, and are unreliable at tasks without ground truth. Speeding up progress in any of these areas seems likely to speed up general AI R&D, giving everyone else less time to execute time-bottlenecked agendas (e.g., trying to do Plan A).
The second reason I don’t feel good about this is because I’m less confident it will help make handoff/deference/superalignment go well. To hand off conceptual alignment research to AIs we need to trust them to 1. be good at this research and 2. be generally trustworthy/aligned. We still don’t know how to reliably prevent prosaic outer misalignment issues (e.g., sycophancy or going off-constitution), let alone worse issues that will make AIs less trustworthy in the coming years. For this reason, I expect #2 to be more of a bottleneck to high-stakes alignment research than #1, which seems to be more likely to be an emergent property of more capable models.
I haven’t seen anyone clearly write up these arguments and I think that more public dialog on dual-use AI safety work would be beneficial. In this post I:
Give examples of arguments people have made for the acceleration of alignment-relevant capabilities
Give arguments for why accelerating capabilities differentially useful for doing alignment research shouldn’t be a current priority
Address two counterarguments:
“But what about things like philosophy?”
“But isn’t it pretty unlikely that you help labs make real capabilities progress?”
Examples of arguments for the acceleration of alignment-relevant capabilities
I’m including these examples to give a sense of the kinds of claims being made and what motivates them, not to be comprehensive or fully explain each argument.
Important disclaimer: Note that the rest of the post should not be read as responding to any of these claims specifically (many/most of my arguments don’t apply to all of them) but instead as responding to the general family of arguments shaped like this.
Example 1: Quote from “How do we (more) safely defer to AIs?”:
Although I personally think this is a bad idea, I find it laudable that Greenblatt and Stastny point out the potential tradeoff explicitly. At a high level, they argue that for AIs to be good at alignment research, they need to be broadly aligned, good at AI R&D, and good at messy conceptual things. They also argue that handing off safety research is safer in less capable models, hence the suggestion to differentially accelerate the skills needed for alignment research without pushing general capabilities too much.[1]
Example 2: Wei Dai in “Increasing AI Strategic Competence as a Safety Approach”:
Dai’s argument explicitly states that alignment is a prerequisite for this to work which I agree with. He also explicitly flags that improving strategic competence could make things worse:
Although Dai wants AIs that are good at arranging a pause rather than solving alignment, I argue that it’s analogous to the Greenblatt et Stastny approach. The idea is still that there exists some capability that could be differentially good for safety and we should consider accelerating it.
Example 3: Quote from an update on Manifund for a project housed at Redwood Research:
As I understand it, the idea is by making AIs better at conceptual reasoning they will be able to help better at acausal stuff and maybe a wide variety of things that could prevent human extinction or other catastrophe also.
Other examples:
Superhuman Articulacy as an LLM Safety Target
Paul Christiano expresses that advances in agent capabilities could be positive in older writing here but does not say that people should work on this.
Why we should not do this kind of differential acceleration
Alignment bottlenecks are also capabilities bottlenecks
There are certain problems with current AI systems which make them bad at AI R&D:
They perform worse on “messier” tasks compared to those that are cleanly verifiable
AIs are limited by ‘creativity’ or ‘research taste.’
There are additional problems that make them bad at alignment research:
They can’t reason in domains without ground truth
They are bad at philosophy
Rereading these two lists, they appear to rhyme with each other. To be fair, alignment research does, at a glance, look like the kind of thing that would require more conceptual / philosophical breakthroughs compared to capabilities research. So there are good reasons to think there exist some philosophical skills which could be differentially useful for alignment.
But accelerating reasoning without ground truth broadly seems too bottleneck-y for both to be a good safety target. For other capabilities that appear more narrow, there is the risk of unfavorable generalization. For example, it’s likely that being good at messy reasoning for a somewhat narrow skill requires eliciting more general conceptual skills that apply to other domains. We have seen some examples of this empirically: training models on logic problems can induce new reasoning abilities and generalize to unrelated math benchmarks. Likewise, the “let’s think step by step” technique generalized broadly and motivated new training techniques even though, on the surface, it is simply an elicitation technique for deeper reflection.
At the very least, there is overlap between alignment-capabilities and capabilities-capabilities and, in expectation, some speedup in AI progress conditional on successful differential acceleration.[2] I don’t think it will be trivial to minimize this risk because various stages of the research can be infohazard-rich and even vague details about techniques can be enough to reproduce them.
This could hurt time-bottlenecked strategies
Old-school EAs used to talk about being a force-multiplier. For example, instead of directly distributing bed nets or something you could trade stocks, make millions, then pay the salaries of 10 people who distribute bed nets or work for the bed nets charity. If you simply worked at the charity, you would be only responsible for 1/10th of the bed net impact so you have effectively multiplied yourself.
In AI safety, anything that accelerates AI progress and brings us closer to RSI is effectively a force-minimizer. Shortening timelines gives a large number of projects less time to figure stuff out and lowers the probability that any of them are successful, effectively subtracting workers from those efforts. There are some agendas where the progress on safety-relevant capabilities would speed things up but there are some activities which are bottlenecked by time. Concretely, making timelines shorter means there is less time to advocate for a pause, fundraise for a promising new alignment technique, develop compute verification technology, etc.
Importantly, I would not expect any capabilities acceleration to be differentially useful for things like advocating for a pause. Imagine a world where we have AIs that have the skills that would make them good at lobbying for a pause: because labs will have access to the most compute and the best capabilities (they may not make the best models public), they will have the advantage to advocate for the position they want (which would not be a pause probably).
I don’t think that this would unblock superalignment/hand-off plans
(This is all written under the assumption that superalignment-flavored strategies are a good idea which they may not be. A lot of the time when people say “differentially good for safety” they mean “differentially good for superalignment/handoff” but here I am assuming that this is fine.)
If a future AI system were capable of good alignment research would we trust it to do a good job? I would argue probably not: current AIs can already be sycophantic, dishonest, pretty misaligned, and can go off-persona or off-constitution in undesirable ways. We can expect more outer misalignment failures as we scale up increasingly insane post-training schemes and even possibly inner misalignment problems at sufficient levels of capabilities. In other words, it’s not clear we are on track to have a model that we can trust to hand off our safety research to.
Concretely: working on alignment directly seems just as differentially good for superalignment/handoff compared to any kind of capabilities acceleration even if we are ignoring the capabilities externalities entirely. (Of course, alignment work can be dual use as well,[3] and for that reason I think it is reasonable to be concerned about alignment research that doesn’t aim to solve the core alignment problem but rather makes incremental progress on product alignment.[4]) The argument for alignment becomes stronger once we account for the externalities and other factors:
Alignment progress may be useful for a wide variety other alignment agendas.
Alignment work has limited capabilities externalities compared to trying to accelerate certain safety-relevant capabilities.
The capabilities bottleneck will likely be solved by the time it matters most: either by default (e.g, with scale you start to get this) or by something more intentional (e.g., downstream of labs trying to train in research taste). Alignment by default is also possible, but conditional on it being true, its unclear if any of this matters because we may just be in a good world.
But what about things like philosophy?
I grant that more narrow capabilities like making models good at philosophy are less likely to be useful for automating AI R&D and probably pretty useful for alignment. But is it possible to develop techniques that only push philosophy and nothing else? If you train on messy philosophical thinking or you get good at eliciting it, would you also be good at more general messy conceptual thinking? Is it even possible to be good at philosophy without being generally really good at messy conceptual thinking? Likewise, if the elicitation technique truly only targets philosophy, could other actors not use a similar technique to elicit more research-relevant conceptual thinking?
Overall, I’m not convinced that narrow capabilities can just be selectively advanced through elicitation or other techniques (or that the same technique could not be applied to additional narrow capabilities).
But isn’t it pretty unlikely that you help labs make real capabilities progress?
(It’s somewhat rare I hear someone make this argument but I have heard it a few times so I thought I would address it.)
I find people can be pretty confident that accelerating safety relevant capabilities like messy conceptual reasoning is really quite tractable and labs are ignoring it. But then when talking about the capabilities externalities or how this could speed up timelines, there is a feeling that pushing the frontier of AI agents is really quite difficult because there is already billions and soon-to-be trillions invested in this and all the low hanging fruit has been picked.
I’m open to the idea that trying to advance capabilities relevant for alignment is intractable and the result of efforts here will be unsuccessful and therefore net-neutral (ignoring the opportunity cost of the money/time spent on the research). But I think it’s hard to hold 1. the research is very tractable, 2. there is overlap with capabilities and 3. any low hanging fruit with capabilities research has already been found. I suppose you could reject #2 but as discussed above, I don’t think we can cleanly split “safety-capabilities” and “capabilities-capabilities.”
I will also note that the open-source frontier may be easier to advance and pretty dangerous also.
My epistemic status
Thinking about superalignment or what actions in the future make it go good/bad requires a certain amount of playing 4D chess with a blindfold on. I think people should be uncertain. I myself am uncertain that superalignment-flavored things are a good idea to begin with. I also don’t think I should be entirely confident that, for example, alignment of AIs capable of alignment research will be a bottleneck to handoff because who knows. The world will be weird.
When I say “differential acceleration of alignment-relevant capabilities is a bad bet” I don’t mean I’m confident it’s negative EV. I’m more saying that this particular bet looks like it takes on a lot of unnecessary risk when there are other things that look equally promising. If the goal is to hand off to AI early-ish (which I’m not claiming is good or bad) then even some prosaic alignment things like trying to systematically understand midtraining or trying to “solve” eval awareness seem less risky while being equally productive and maybe more tractable.
I will say that I feel pretty confident that the dual-use aspects of this kind of research are a real risk that deserves to be more widely discussed and publicly debated. At the very least, it seems good for those pursuing dual-use research to subject themselves to lots of red-teaming (if the idea itself isn’t an infohazard, this should be public) and have a preregistered policy for handling potentially dangerous information.
Greenblatt and Stastny also have ways to address some of the points in my post, but I don’t respond to these specific points directly in an effort to keep the scope broad. (I don’t want to respond to a specific post but instead give reasons for why I don’t find the family of arguments convincing.)
I would argue that even elicitation techniques can be silently risky. For example, “let’s think step by step” seems like somewhat benign elicitation technique that could make the model better at messy philosophical questions. But it also led to reasoning models which was a breakthrough.
I worry that other elicitation techniques could be more broadly applicable than people think because there may be a more general mechanism behind a seemingly narrow capabilities breakthrough.
This is a crux though. If people gave successful examples of narrow capabilities being accelerated in language models where there is no more general mechanism that could be used for other domains, I would update my beliefs. I tend to think that if there is an elicitation technique that is good at surfacing safety-relevant capabilities, there is a good chance it could be repurposed for surfacing arbitrary capabilities.
Working on this misalignment bottleneck can mean a lot of different things and some of those things also have capabilities externalities so we can run into similar problems. RLHF is a standard example of “try to solve outer alignment and accidentally improve general capabilities.”
Things like trying to make the model less eval aware during safety-relevant tests seems robustly good with limited negative externalities. Things like trying to make AIs less reward hacky also falls under the umbrella of “solve outer alignment in early transformative AIs” but I can imagine this making general AI research agents easier to use for non-safety things.
This is all to say that dual-use risks should remain a consideration even when working on things that appear more alignment-flavored.
It’s best to try to solve the core problem but admittedly most alignment research does something closer to “incremental progress to make the alignment metric go up for some model organism on some eval.” I actually generally cautiously support the latter in addition to the former. If we don’t “solve” alignment then we are in a world where we are putting lots of different bandaids on the problem and hoping for the best. This is extremely reckless and people working on bandaid development should understand this and communicate this honestly, but if we are in this world I think we will be pretty grateful that there are different bandaids to choose from.