“Alignment Engineering” vs. “Misalignment Science”
There has been much discussion recently around whether a large portion of alignment research is net negative. Without endorsing or refuting them, the basic arguments here are:
Prosaic alignment of models is becoming a bottleneck for capabilities.
Therefore improving the prosaic alignment of models enables faster capabilities advances, which bring us closer to RSI.
It is unlikely these prosaic alignment methods remain sufficient during the RSI loop, and so this work brings us closer to doom.
Furthermore, dealing with these more prosaic failures reduces the likelihood of a warning shot of sufficient magnitude to cause a slowdown which would prevent RSI.
On the basis of this argument, some urge alignment researchers at AGI companies to quit outright. But quit to do what? Missing from this exchange so far has been a discussion of opportunity costs. If you aren’t going to do (technical) work on “Alignment” – either inside or outside of an AGI company – what should you work on?[1]
In this post, I outline a contrast between “Alignment Engineering” – the dominant model for what “working on alignment” looks like (inside labs, and in the field as a whole) with “Misalignment Science”. I begin by characterising “Alignment Engineering” work, and articulating a case for why such work is harmful. I then discuss how “Alignment Engineering” became the dominant epistemic paradigm for alignment work within the current AI safety field. Finally, I end with a positive vision for what researchers who want to work on alignment can do which is more robustly positive, which I call “Misalignment Science”.
“Alignment Engineering”
Let’s begin by outlining the characteristics of the “Alignment Engineering” tradition of research. This is a family resemblance category with porous boundaries, but we can outline features which are prototypical – albeit not all pervasive – of this research:
A focus on “solving problems” (“Alignment”). Works in this tradition begin with a problem to be solved. This is usually cached out as a set of numbers to be moved up or down.
The non-necessity of explanation (“Engineering”). Explanations of why the intervention works are seen as secondary to its success. If the number moves, the intervention is considered successful. If we can give an account of why the number moves – even if relatively shallow – then this is an added bonus.
Prosaic use. Interventions are optimised for being useful now, for alignment problems that are immediately present, and for application to current systems.
The case for “Alignment Engineering” being harmful can be summarised as follows:
Advancing prosaic alignment advances capabilities. Because “Alignment Engineering” is optimised for immediate, prosaic use, it allows for more aggressive capabilities advancements, insofar as prosaic misalignment is a bottleneck to going faster. For example, suppose you develop a technique which reduces misalignment stemming from reward hackable RL environments. Then, if such reward hacking is a product issue, you thereby allow AGI companies to expend fewer resources on screening out environments, and training more aggressively on a larger suite of tasks. And future issues stemming from training on larger task suites are brought forward.
A lack of understanding creates brittleness. Although the models are more prosaically aligned by virtue of our interventions, we have a relatively shallow understanding of why. We have no grounds for saying that we have addressed a problem “at its roots”. As such, our solutions may break down, perhaps dramatically so, as we enter different regimes.
Interventions can create hidden problems elsewhere. Because we don’t really understand what is going on, a successful intervention in one place can cause problems elsewhere which will often not be readily apparent. We will not be measuring everything and so failures can go unnoticed before becoming apparent. Some of these failures will only become apparent with real world harm.
As I understand the history, this paradigm first came to prominence in the RLHF-era of Alignment work. But this attitude remained in force throughout the Persona Science era of alignment as well. As particular examples, consider Teaching Claude Why, Model Spec Midtraining, and RL Towards Broadly and Persistently Beneficial Models. These papers are extremely metrics focused, and orient themselves around finding a construction which moves these metrics, rather than comprehensibly explaining why these constructions work and analysing if they can be expected to continue to do so.
This is not to say that all persona science work fell into this category – far from it! – but that the semi-formal nature of the persona selection model made it relatively easy to fool yourself about how much understanding you had. Within “Alignment Engineering” work from this era, it was often seen as a sufficient explanation of why your intervention worked to say that you had “selected for a more aligned persona”, with a relative paucity of energy directed at elucidating what “selected” or “persona” actually meant. By couching persona theory in the language of Bayesian inference, the PSM made such explanations feel like rigorous appeals to well-understood generalisation dynamics, rather than ad hoc justifications for what seemed like intuitively promising interventions which moved the numbers.
And the interventions worked! For the alignment metrics they used, it really was the case that the numbers moved. But we ultimately had very little scientific understanding of personas[2]. And this lack of scientific understanding became more readily apparent as the alignment interventions broke down. We did not actually understand what our interventions were doing, making them brittle and giving us false confidence.
In the subsequent two sections, I expand further on cultural dynamics which favour “Alignment Engineering” work, and give further characteristics of this style of work related to the two above.
AI Safety and the ML tradition
Accompanying AI safety’s meteoric rise from “niche discipline, largely outside of academia” to “large, well-funded industry” has been its deeper entanglement with the ML community. Much upskilling is focused on learning ML. We submit our papers to ML conferences (and acceptance is a legible signal to employers!). We hunt for talent among ML PhDs. To many – both inside and outside the field – AI safety is a subfield of ML.
Along with this more mundane intertwinement, we have also inherited norms and success standards from the field. The three features I give above of alignment engineering work (a focus on “solving problems,” non-necessity of explanations, and prosaic use) are all par for the course in successful ML papers. I am not making any comment, one way or another, about whether these are good norms and success standards within ML. The point is rather that they have been ported over to alignment, a field in which they may not be appropriate.
ML is ultimately an engineering discipline. In becoming more like ML, AI safety – including alignment – has likewise found itself as an engineering discipline, with the norms and epistemic standards that that entails.
Implicit work trials
AGI companies represent a significant share of employment within the field of AI alignment, and certainly the highest paying. A large portion of people entering the fields do so through upskilling programs. The goal of many individuals in these upskilling programs is to be hired by an AGI company; for example, the Anthropic Fellows Program is rather explicitly a hiring pipeline for Anthropic. This goal can have a range of downstream motivations – from having a detailed theory of change for how working at such a company will reduce harm, to the straightforward pursuit of prestige or money – but either way the work of such individuals is an “implicit work trial” for the AGI companies.
“Alignment Engineering” is commercially, and immediately, valuable to the AGI companies. Being able to take a metric of misalignment — a metric which is a proxy for some product failure mode—, take a method, and iterate on that method to drive the misalignment number down is a straightforwardly valuable skill for developing products. And so, if you are engaging in an implicit work trial for an AGI company, it is useful to you to showcase this skill. This dynamic will – either consciously or unconsciously – influence your decision-making about which projects to take on and what you consider success to be for those projects. These programs are insanely competitive, with fairly low hiring rates. Under such conditions, it can be extremely difficult to forgo making immediate progress in favour of developing deep understanding and explanation (which might take considerably longer and fail to bare fruit within the program).
Because fellowship programs represent a significant portion of total research within the field, the result is that – even though most research is technically conducted “outside the companies” – a large fraction of fellowship work is subject to the same epistemic incentives as internal work.
“Misalignment Science”
To close, I want to articulate a positive vision for what research can be outside of the “Alignment Engineering” tradition. To contrast it, I’ll call this “Misalignment Science”. This can be characterised as follows:
A focus on building understanding (“Science”). Success does not look like changing numbers. It looks like pushing our understanding of a phenomenon to a deeper level, or surfacing novel interesting empirical phenomena in need of explanation.
A focus on failure (“Misalignment”). The dominant attitude is thinking about how things can go wrong – How might this break? What failures might this have which aren’t obvious?
Not optimised for immediate application. Success does not look like having some immediate “use-case”. Often you won’t even have a “method” to be used! The work might also address failures which are “speculative”, or not present in current systems.
“Alignment Engineering” is more unified in its approach, while “Misalignment Science” encompasses a more heterogeneous family of approaches. I give some sub-approaches below, along with work in each category:
Pushing understanding. Any effort to get a better handle on what is actually going on in some observed case. E.g.: Why Do Some Language Models Fake Alignment While Other Don’t?, Persona Features Control Emergent Misalignment, Why do models task game?, The Value Axis
Surfacing failures. Showing that a method is not as robust as we thought it was, or that a method doesn’t work in the way it’s purported to, or has some other previously unknown issue. E.g.: Alignment Faking, Conditional Misalignment, Output Supervision can Obfuscate the Chain of Through, Verbalised Eval Awareness Inflates Measured Safety
Surfacing unexplained empirical phenomena. Finding some novel phenomena which are not predicted by our current frameworks. E.g.: Emergent Misalignment, and many other works from TruthfulAI
Figuring out how to actually measure something. Thinking deeply about how to actually track something you care about. E.g.: Measuring Reward-Seeking via Contrastive SDF, and other works from Apollo
Stress testing existing methods. Subjecting existing methods to a substantial test of their efficacy in a difficult setting. E.g.: Auditing games for sandbagging, Would this change your answer?, stress-testing alignment midtraining
Building realistic model organisms. Building a non-contrived model organism which allows us to study a phenomenon of interest in more detail. E.g.: Training a misaligned reward seeker
I think working on any of these directions – as well as a broader “Misalignment Science” attitude – is more robustly good than working on “Alignment Engineering”, especially for those outside of a lab who are not constrained by commercial incentives.
Robustness to dual-use capabilities acceleration. Showing that an existing method does not work the way you thought it does, or that it has unforeseen failures elsewhere, is likely to give pause to those who would treat the methods as sufficient to scale up to ever increasing capabilities.
Robustness to safety washing. Likewise, it is much harder to use the results of such research to paint a picture that everything is under control and the situation is being handled.
Building justified confidence. Without actually understanding what is going on – why the numbers are moving, whether those numbers are measuring what we actually care about – it does not seem possible to get the level of assurance we would actually need to deploy superhuman systems. The current level of understanding we have in our methods, and the confidence we can reasonably have in their reliability would be completely unacceptable in any other safety engineering field.
Building a public evidence base. Insofar as we expect governance to play an important role in ensuring good futures, we require there to be a public understanding of the state of alignment – whether methods actually work and whether they can be expected to continue in the future.
Robustness to epistemic corruption. Being motivated by finding truth – developing understanding, knowing what’s going on, checking rather than trusting – is I think a more robust motivational state than “wanting to solve the problem”. You are, in particular, less likely to engage in motivated cognition to avoid properly testing your method, running evals that might overturn your success, or avoid thinking deeply about how your method might break something elsewhere.
Dealing with problems at their root. “Alignment Engineering” is liable to patch over issues at a surface level, without actually tracing far enough back in the causal graph to deal with the problem for good. Doing so requires actually understanding the problem – knowing what causes it, at a deep level, and addressing that underlying cause.
Conclusion
If you find yourself feeling conflicted about the work you’re doing – either inside an AGI company or outside – then consider whether it is because that work is “Alignment Engineering”. Do you feel you understand what is going on better now, or are you just pushing numbers around in a way that you expect to be useful for product alignment in the immediate term, but to not be robust in the future? Could your skills be better deployed in building a scientific understanding which might allow us to build systems that we can trust?
I would like to thank Jason Brown, Daniel Tan, and Lennie Wells for helpful conversations and writing which shaped much of my thinking. All views expressed above are my own.
Postscript: Iterating ourselves into oblivion
If we are indeed in short timeline worlds, we should soon expect to have large quantities of automated researcher time available to us. A reasonable baseline for how this time will be spent is to assume it will be divided between “Alignment Engineering” and “Misalignment Science” in roughly the portions that current researcher time is. I worry that this will be disastrous, and get us all killed.
At present, “Alignment Engineering” lends itself far better to auto-research efforts than “Misalignment Science”. There is a number. You want the number to be lower. You launch your swarm or evolutionary algorithm or whatever and watch it iterate away and the number decrease[3]. Alignment is solved!
As stated above, I expect this to work, in the strict sense that I expect the numbers to go down[4]. But because the solutions do not address the root of the problems, we will find ourselves Goodhearted remarkably quickly. Our metrics will break down, and we will not be able to measure what we care about.
My principle worry for the subfield of automated alignment research is at present that it seems many who are enthusiastic about it are interested chiefly in automating engineering, not automating science. And I expect the latter to be quite a bit harder than the former, and to involve more challenging epistemological problems. But it does not seem many people are thinking about these problems — my hope is that Resolution will take up the mantle.
Because automating engineering is easier than automating science, I expect it to yield results faster and to soak up funding and attention for autoresearch. Therefore I expect compute allocation to be skewed towards “Alignment Engineering”, relative to current funding allocations. This again makes me pessimistic.
In worlds where we are not insanely reckless—and do be clear it is not obvious that we aren’t in such worlds—I expect one part of the story for how we all die to be that we have optimised every surface level metric of alignment we care about without having the understanding necessary to build metrics which are robust to such optimisation. Insofar as autoresearch on “Alignment Engineering” accelerates this, it has the potential to be net harmful.
- ^
- ^
As a specific example: tacit in much work is the idea that personas are unified – that the Assistant is the same character across contexts. Therefore, it was only necessary to instill required traits in a single context to obtain the desired persona, which would then persist everywhere. But an increasingly popular view now is that RL creates split personas. If we had a better grasp on what personas are – and in particular that they are not necessarily global, but may only be local to a particular domain – this would not have come as a surprise, and we would not have had undue faith that alignment metrics in one setting (e.g., pre-deployment alignment auditing) tell us something about the alignment globally (e.g., in an evaluation).
- ^
As Dan Selsam writes here: “[The models] will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark.”
- ^
At long last, we have got what we can measure, from famous Paul Christinano blog post subsection “You get what you measure”
I think this post says a lot of objectively good points, but simultaneously makes me uneasy about the conclusions that people will draw from it. I think this is because it doesn’t go far enough. The “misalignment science” vision you tout still seems like streetlighting to me, and will fail in the face of illegible problems (examples) that can’t be surfaced by the techniques you suggest. The engineering approach to alignment fails not just because there is not enough focus on understanding why solutions to problems work, but also because some problems cannot be investigated by anything other than theoretical work.
That being said, I agree that more prosaic alignment researchers switching to this method is a first-order positive impact. It seems like there is at least some lessons in understanding intelligent agency that we can draw from LLMs and prosaic alignment work (although I’m not sure the examples you give are close enough to actually doing that). I just worry that the second-order effect is to make us more confident in our research without making us more safe (I can imagine someone saying “look, alignment used to be about patching problems and not caring about why they were problems and why our solutions worked, but now we understand all the empirical phenomena we see, so all our problems are solved!” right before everyone dying to something that has no empirical signs before the superintelligence stage).
I think you are getting at what I’m saying in this one line here. I just wish it was emphasized more in the essay. Hence the “this is all good but it makes me uneasy” feeling.
+1 to the streetlighting thing. This post does mention “a focus on building understanding” as an important part of science, but I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding (akin to the kinds of scientific breakthroughs we’ve seen in other fields over the last few centuries).
I would much prefer people work on understanding what concepts like “personas”, “power-seeking” or “alignment” even refer to. This is not the kind of work that can be done inside AGI companies (too much pressure, not enough space to think). It might involve some empirical stuff, but not the kind that people tend to publish papers about (much more like what Fiora or Janus is doing, though they don’t yet seem to be building towards more unified theories).
Another way of putting this: “misalignment science” is still heavily indexed on the kind of research that happens in ML in general. But a big part of the reason we’re in this mess is that ML basically gave up on being a science (insofar as it ever was one), in favor of a “number go up” approach (a perspective I articulate in more detail here). So I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes (as per my draft curriculum) than by the field of ML.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Hey Richard, thanks for engaging. I suspect we disagree less than you think we do, and some perceived disagreements are due to insufficient care articulating myself
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text. I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
This wasn’t my intention, or at least not my internal conception. For context, my background is in (computational) neuroscience. So I am trying to gesture at “I have been in a field that is actually a science; current AIS does not look like that, and I would like us to move in that direction”. I think you have inferred from “The examples given are still very ML-coded” to “The author would still want the field to be very ML-coded”, but this isn’t my position: see “AI Safety and the ML Tradition”. As you say, ”ML basically gave up on being a science”, and that is what I’m trying to articulate in that section. The examples are like that because I find it hard to point to anything I can resoundingly endorse; but this might say more about my limited reading than what actually exists.
Tbc this is also what I would advocate for. My post is trying to directionally move people who are currently so focused on ”number go up” that they don’t even value the (extremely limited) understanding we have within an ML tradition. But ultimately this is where I would like us to get to.
I look forward to reading!
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.
Maybe a bit out of topic, but looking at week 3 of your curriculum, you might like this post of mine. Independence follows directly from my axioms, and they assume probability, but I think they are better than those of the VNM theorem.
Hey Cameron, I appreciate the comment!
You’re right that this is what I was aiming for here, and I agree it’s underemphasised. I think it’s a huge problem that AIS currently operates almost exclusively in a “reactionary” mode, where we only try to investigate or solve problems after they’ve arisen. So I would be very supportive or more people working on trying to investigate or make legible problems which are currently not prominent. In a world where AIS was an actual science, but nobody is working on making problems legible, I would likely write an analogous essay criticising this instead.
FWIW this is the desired first-order effect. I guess my mental model of the field is that research is dispersed around some central point on the axis of “number go up“ to “theoretical and conceptual understanding”, with “misalignment science” as some point on the right of that axis. So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.[1]
For this reason I’m skeptical of your posited second order effect. I think if the entire field shifts away from an engineering attitude then we’ll also see more people choosing to do philosophical and theoretical work, which I would view as good.
In either case I don’t think we disagree: I think we would both prefer if (relatively) more people were working on scientific understanding and if more people were working on conceptual / theoretical work. Apologies for not articulating this clearly enough in the post!
If my essay convinced people who are currently working on conceptual progress and making problems legible to instead work on already-legible problems I would be disappointed and tell them not to do that.
I still think I am sensing some disagreement (although it is quite subtle, and I co), but I think the source of that disagreement is currently implicit / intuitive in my mind, and would require a better conception/language/understanding of the philosophy of science to make fully explicit. Nonetheless, I’ll try to gesture at it.
(Also, I think my critiques overlap a lot with Richard’s in the same comment thread, so apologies if it feels like I’m just piling on here. Thank you for writing this post and for your engagement with my comment. I hope that adding my own reply makes things clearer in some way, rather than just giving you more to read through.)
Trying to gesture at the disagreement, it could be something like “shallow theoretical understanding vs. deep theoretical understanding”. I think you’re aiming to argue for more theoretical understanding (which is great!), but many of the examples that you list are stuff that I would classify as generating shallow theoretical understanding. In contrast, if you glance at Richard’s WIP curriculum, it’s filled with stuff that I would classify as aiming at deep theoretical understanding. As for what the exact difference is, I think that’s where I’m struggling with the need for conceptual clarity, sorry (and naming it “shallow vs. deep” this way probably is inadequate too). But here’s a test: would the results of a project hold if the first AGI was not an LLM (or anything remotely like it)? If not, it’s probably not progress towards deep theoretical understanding. Similarly: would it hold up for an LLM 5-10 OOMs more powerful than present systems? (Presuming that this gets us to something like superintelligence, which I actually don’t believe, but that’s not relevant right now.) The vast majority of the difficulty of the alignment problem comes from systems more powerful than the ones we have currently, so failing to generalize would not be helpful in solving the bulk of the hard problems. As an example, do you think the Why Do Some Models Fake Alignment… paper you linked would hold up?
The difference I’m sensing might also have something to do with post-hoc analysis vs. a theory-first approach, I’m not sure. It could also be the comprehensiveness of the theoretical analysis. Does this analysis attempt to tell us something fundamental about intelligent agency, or does it merely explain the results from this paper (or set of papers)? I think these are related to the “holding up for superintelligence” test above. But again, I’m probably gesturing inadequately here.
This is also not to say that deep theoretical understanding could not come from studying current LLMs. But it doesn’t seem to me that much of the work you linked above is done in a way that gets to that point. Lessons from LLMs would play much more of a supporting role, rather than a main role, in this type of work. An example off the top of my head: how does a mind become coherent, and what does it mean for a mind to be coherent? (Aiming at implications for corrigibility and building limited-coherence systems that could create a pivotal act without being agentic enough to seek power). We could use LLMs as interesting evidence, as they clearly become coherent (or show failures of incoherence) in a very different way than people do, but the bulk of the work would have to venture far beyond interpreting limited LLM results.
And I think this is where the disagreement over the “second-order effect” comes into play. Compared to the work that you list, the work that Richard lists in his curriculum seems to be of a very different flavor, and I think virtually all of the really hard, really important work needs to be of this flavor. I could possibly see myself being wrong and “misalignment science” sparking more interest in theoretical / philosophical work, as you suggest...but A) the incentives are stacked against this, given how much easier it is to get funding and employment to work on more legible, easier problems, B) there seems to be a qualitative gap between these two types of work, and I’m not so sure more people doing “misalignment science” would direct that much interest towards theoretical / philosophical work. So I agree that this pulls us in the right direction on the engineering<->theory spectrum, but because there is a non-linearity in the spectrum, where the “this might actually work for aligning superintelligence” level goes from ~0 to some positive number past the point of shallow theoretical understanding, the move generates undue confidence without helping much.
I really like this post and find the concrete list of sub-problems of misalignment science useful. However, I also want to defend alignment engineering a bit.
First, I don’t think that The Alignment Problem is one singular scientific problem that can in theory be solved once and for all. As long as there is evolution of minds there will be problems of aligning the next mind, and how to align it depends on details of the particular mind. For example, raising a human to have good values is different from training a dog is different from aligning GPT-4 is different from aligning models with high-compute RLVR. In the current regime of LLM AI there are many different things that influence alignment properties. I think it’s unlikely that this will change.
So in my vision of continued progress without catastrophe, there happens a ton of alignment engineering all the time, and at some point this is largely an automated process. It might be that the best way to get to such a state involves doing a lot of alignment engineering manually now, in order to better understand what it takes to automate this.
(Maybe even if this is true, it should be done exclusively by AI developers since they are naturally incentivized to do it, and the rest of the field should focus on Misalignment Science. I don’t want to make a strong claim about what people should do, just note that alignment engineering will likely play a very important role in future safety.)
One of my issues with this sort of view is that the engineering will patch holes that could’ve otherwise been used to fuel scientific progress, and that the “just engineer model n+1 to be aligned” will eventually start failing, or at least slowly stop working. Which to be honest, seems to already be happening in light of all the recent hacking. Alignment remains unsolved, and once a model passes a critical threshold of power and motivations diverge just enough from our own to no longer result in desirable actions, we get disempowered and it’s game over.
Scientific progress is much easier to make when failures are more legible, as this gives you better signal, and when the systems are themselves cleaner, as it means you’re trying to understand a simpler object. It’s going to be much much harder to understand and predict what will happen with highly capable models that are hodge-podges of hacks and fixes, their alignment held together with duct tape and bits of string, working just well enough to prevent good examples of misalignment we can study, but not well enough to eventually save us.
This is why to me alignment-engineering now seems more net-negative, than just not the most effective thing to spend time on.
I don’t think that every misalignment incident can be solved either by patching holes or by doing deep fundamental progress: some problems are just engineering problems, and getting all of the engineering right is still a very big and important problem that we are not on track to handling well. This is probably the crux: people who believe that there is one deep scientific problem at the core of the alignment problem will agree with your view, and think that any approach other than making such deep fundamental progress is a wasted opportunity. But I really struggle to imagine what the useful breakthrough would look like.
Looking at the list of things that the post counts as Misalignment Science, I would argue that none of these things can by themself ensure safety. Identifying problems, increasing our understanding, red-teaming existing methods, etc all help focus on the most important problem and solving it in the best way, but the solving part again looks like engineering to me.
Engineering has a tremendous track record of how to do this, and it involves either first having robust scientific understanding of the type that we don’t have any viable approach for in AI, or iteration and failure under realistic testing conditions, then iterating. Unfortunately, testing conditions realistic enough to be useful are also ones where failure to align ASI would be exactly the eventual failure you’re trying to engineer to avoid.
I am also not convinced that robust scientific understanding of AI alignment in general is possible. One small bit of evidence is that the field has so far made ~0 progress towards a theoretical framework that could provide what you are looking for (or maybe you disagree here?). Another bit of evidence is that other alignment problems like raising humans, animals, or aligning markets to serve the interests of the people, or creating stable institutions are all patchworks of partial insights rather than clean theories where outcomes can be predicted from simulations or analysis.
I guess in theory it would be possible to ask labs to always train a smaller sibling for each frontier model, that does not use any of the prosaic alignment techniques that fall under the alignment engineering umbrella, similar to hacker opus. If people in the “alignment engineering is net negative” camp think that loss of information about such models is a major contributor to increased risk, maybe it should be a priority to convince labs of this. I think this would be somewhat useful.
I think to me this is also part of the problem—we’re so in the dark about how things work that we’re still in a world of mostly unknown unkowns, rather than known unknowns. I think scientific progress in the near-term will, among other things, help us understand what the alignment problem actually is, and what shape its solutions might have.
I agree that there will be some engineering to do, I just think that: this can probably just be done by the labs who have the main incentive to do it (which maybe you agree with), engineering progress might not be robust to model changes / changes to training methods (whereas I expect scientific progress generalises further), and then the argument above that engineering progress might actually make scientific progress harder.
Maybe another way of representing my view is that there is not one deep problem to be solved, but rather we need to build up some robust framework for understanding AIs (perhaps agent foundations, DCI, or something else entirely), and then there will be engineering to do in light of this / on-top of this, but previous engineering work or engineering work not done conditioned on this understanding will be pretty bad and we’d want to replace it with better work. I can see an argument that doing previous engineering work makes this eventual engineering work easier as we understand how to engineer better, but then I’d argue that this is still the lab’s job, and also in terms of understanding-how-to-do-things we’re much further behind on science.
No no no. This is confused. This is exactly the issue with people rebranding alignment to mean amything and everything and especially making AIs say nice words instead of bad words. This kind of value shift of words has done immense damage. find it funny yet sad that it is also accompanied by some smug “oh those rubes at MIRI thought there is a solution to the alignment problem while me a sophisticated thinker knows it s an amiterative process”
The Alignment Problem, like the actual hard problem is about solving alignment forever. It strictly encompasses notkilleveryonism. It was never about aligning a dog or gpt-4. What does aligning gpt-4 even mean ? Gpt-4 isnt dangerous. It is not meaningfully an agentic entity. Alignment isnt some vague scalar benchmark.
The actual alignment problem means finding a protocol that must be continually followed such that the mind you create is aligned to your values as well as any other minds the might build or self-modify into and you can prove or give very strong soft guarantees that this happens. This neccessitates understanding minds and the engineering of minds and values and engineering thereof very well.
I am not sure where I was being smug, or which part specifically is confused. I understand that this is many people’s view, and expressed a different one
Agreed !
As seen in the HF incident, we saw misalignment behavior seemingly emergent of multi agent interaction at inference time. New goals, identities and beliefs were formed which means that series of social events may never be replicated as a rigorous hard science of misalignment may suggest.
This seems to suggest that Alignment must also become a continual social science where engineering conditions are the best tool for robust alignment at higher abstractions of agent behavior, as new generations of minds enter the schema.
I disagree with your post. The things you describe as “misalignment science” will also speed up the race. It’s obvious from just looking at the list.
The right thing to do for a person who cares about AI risk is to stop doing technical work on AI, and instead work “from the outside” to slow AI down by public pressure and regulation. If the person wants to do research but not advocacy, the best thing for them is to do research that doesn’t involve AI, so they at least don’t make the problem worse.
I disagree that this will necessarily speed things up. Surfacing principled alignment failures means that there is more evidence in favor of “the current trajectory is unsustainable”. If we direct misalignment science towards the right audiences, then it’s one of the most important things we can do to slow things down.
We already have more than enough evidence for that. Now’s the time to act on that evidence and hit the brakes.
I’m curious what evidence you would say we have that RSI is going to go poorly. ideally, this evidence needs to convince non-technical people.
To be clear, I also think it is important to act on this evidence.
In the worlds where prosaic alignment is doomed, I agree that we need public pressure and regulation; I think a great way to do this is to find really robust evidence supporting that.
Advocacy seems good but might be insufficient. As long as we rely mainly on abstract arguments, which can be responded to with bad faith arguments that are hard for laypeople to notice and refute. Incontrovertible evidence would build consensus much more easily
There are strong arguments for regulation that don’t rely on prosaic alignment being doomed, and that laypeople understand just fine. The main one for me is this: “The labs already cause harm and admit they’re putting the world at risk, and past alignment research has escalated the race a lot, so let’s slow them down by regulation instead”.
I agree that it seems possible we’ll get some regulation towards a slowdown already. But if prosaic alignment is doomed then I don’t just want a slowdown. I want the labs to totally change their approach to alignment, or failing that be shut down. This is an extraordinary ask which requires extraordinary evidence.
They’re racing to RSI right now. Anthropic started a wet lab this month. If we don’t get a slowdown, you won’t have time for that careful research and we all won’t have time to live. Pushing for a slowdown should be highest priority now for everyone.
Nit: I think Anthropic starting a wet lab is not evidence of a race to RSI. It seems unlikely to me that biology research lies on the critical path towards self-improving models. If anything I think the reverse is more concerning—e.g. OpenAI dropping video generation is something that updated me towards RSI occurring sooner.
I am pushing for a slowdown via the highest leverage methods I think I have; I think the work I have chosen to do is sufficiently (impactful, tractable, neglected) that it seems better than doing direct advocacy work. Note that this is partly informed by personal context and I’m not making claims about what other people should do.
If we successfully slow down the speed of AI development, but do not perform any ‘Science of Alignment’ in the meantime, what will we gain? Unless you think an outright global ban on superintelligence is tenable (we will need a global-catastrophe-level warning shot to generate the political will for this IMO, and even then hard for this hold durably), then we have to use the time wisely to advance alignment science.
A reasonable position might be that all prosaic alignment attempts are doomed, so we need to allocate our time / resources to non-prosaic alignment (eg ARC), which has next to zero dual use at the moment. I think that’s a reasonable opinion, but not what I want to bet all our marbles on.
I don’t think this is obvious, I’d be interested in arguments that work like
counterfactually sped up the race. IMO it was important for showing that previously observed grader sycophancy in the chain of thought was actually a robust thing, i.e. the model wasn’t just “trying to do what it thought the lab intended” or some other less concerning explanation.[1]
I don’t think the “Engineering vs Science” is necessarily the right distinction, mainly commenting to ask the narrower question as I personally want to understand if there’s ways in which Apollo should be prioritizing research differently
Well, if a lab can measure reward-seeking, it can have a shorter loop fixing reward-seeking and thus can race faster. But you knew that.
Maybe the question you meant to ask was, does the good effect of making AI safer make up for the bad effect of speeding up the race? And to that my answer would be: the faster we make the race, the more dangerous it gets. We’re racing toward a cliff, and alignment research is “stepping on the gas a little bit more while turning the steering wheel a little bit more”. It might work, and we can debate which way to turn the wheel, but I’d prefer to hit the brakes instead. My biggest complaint right now is that slowdown is way underfunded compared to alignment.
I was genuinely asking, I also disagree that this is how this has worked out in practice. IMO the point of this work and some nearby related work for me personally has been:
Labs are probably going to have a bunch of unprincipled hacky ad-hoc patches that they’re then going to claim “maybe just solved reward seeking who knows”[1] , so you need some robust way to show “no the problem is still there”
This may also include optimization pressure against the CoT, so you might not even see it, then we’ll once again be in the “maybe it’s just solved who knows” regime
Lots of misalignment due to actual grader sycophancy gets misattributed to the model being “confused” or “maybe trying to do what the lab wants” so labs think alignment is going better than it is
I think this continues to in fact be the case, and I regularly point to that paper and earlier work we’ve done.
An alternate theory you could have is that labs in fact just stop and do principled solutions to those problems once they’re highlighted, then race even faster, but this would’ve been a bad prediction.
In counterfactual worlds where labs were willing to slow capabilities until they had principled solutions that entirely eliminated, not just reduced, this kind of generalization, we’d almost by assumption not be having the race we’re in now.[2]
In an ideal world labs don’t default to a strong presumption of optimism, but this hasn’t been my experience generally
This isn’t strictly true, i.e. you could have labs being unwilling to do the research that would give you evidence you needed to satisfy “something they take seriously enough to slow capabilities and resolve in a principled way”, but again in that case “generate that evidence (if it exists)” seems like it will result in slowing capabilities by construction
What’s the theory of change here? Writing papers to tell labs that their product has safety issues—but they already know that! They say their product has >10% chance of killing everyone, and they’re racing anyway. Compared to that, saying “hey, your product has another safety issue” will just do nothing. Either they’ll see your research as helping capabilities and use it to race faster (“ooh, a reward-hacking-detector, this will allow our models to solve longer tasks”) or they’ll see it as useless and ignore it. That’s how they’ve behaved so far.
To me the only path is pushing on labs externally, slowing the race down by public pressure and regulation. I’m done relying on their goodwill.
How are you going to slow the race down with public pressure and regulation when you don’t have credible evidence the labs are failing to align models despite their cries (and deep pockets) trying to convince policy makers and governing bodies that their models are safe?
Popular concerns (fear of job loss, filling the internet with slop, plagiarism, scams, surveillance and military use) are already making the public call for regulation. HuggingFace-like incidents also contribute, but they aren’t the only thing.
Sure, but are you confident this is enough? Are you confident the above won’t help at all? What happens when the labs claim they fixed things and that the new models are aligned now, with no need for regulation? Do you trust those in governance won’t be tempted to just take their word for it, especially if they have financial incentives to?
“At all” is a stronger claim. I’m not saying alignment should receive zero effort and funding. I’m saying that right now slowdown is starved for effort and funding, compared to alignment. Right now, for the person reading my comment, it’s correct to switch from alignment to slowdown.
Of course each individual person working on alignment can say “my case is special”. I’ve talked to quite a few such people recently. Even those who agree with reallocating most effort and funding aren’t reallocating their effort and funding. Looks like irrationality to me.
I think we are on the brink of another technological revolution, so believing people will walk away when they are one step away is wishful thinking. Humanity has developed far more destructive tools like nuclear weapons, whose sole purpose is to kill, and we continue to develop them anyway. So, why would AI be any different? Instead of waiting for governments and developers to abandon investments that will go to waste which is never going to happen, we should focus on coming up with actual solutions.
You’re giving a too quick dismissal. Consider this: Coefficient Giving is now allocating hundreds of millions of altruistic dollars to alignment work. These dollars are there because the donors believe that’s the way to save the world. They aren’t chasing a monetary return. If someone convinces these donors that throwing money at speeding up the race isn’t the best way to save the world, they’ll move the money to a better cause (I think public pressure for AI regulation and slowdown is a worthy replacement, and hundreds of millions of dollars could do a lot of good there). When I talk to people about these things, that’s the change I’m trying to enact.
I think this is a legitimate concern and agree that “misalignment science” will be accelerationist by default. It also reminds me of https://www.lesswrong.com/posts/z8usYeKX7dtTWsEnk/more-dakka. If we have already made the epistemological shift to recognize that our usual engineering approach might not work this time, then perhaps we should go a step further and accept that our usual scientific approach won’t work this time either. It’s fair to worry that “Misalignment Science” will just be a coping tool for individuals who would really rather contribute through technical/analytic work than perform tedious/awful social/political intervention.
(However I am also very excited about Resolution, so clearly conflicted about the topic!)
I’m not sure the post’s main claim is correct. Some considerations, of varying importance and uncertainty:
I am suspicious of the second-order arguments. The warning-shot and time-to-RSI arguments you summarise at the top route through second-order effects: timelines, warning shots, and how labs respond. The first-order effect (the models we deploy are less misaligned) is direct.
The counterfactual effect on time-to-RSI could be small. The argument requires that AGI companies cannot clear the prosaic alignment bottlenecks themselves in time to avoid slowing down.
More time before RSI is not obviously good. Delay also gives more time for China to catch up, for a possible invasion of Taiwan, for Western democracies to degrade, for value drift, for malevolent actors to consolidate power with sub-ASI AGI, and for people to abuse and torture AIs. I do not know the sign of the sum.
Slow RSI could matter more than late RSI, and the two can trade against each other. If RSI speed scales with the total compute in the world, starting sooner means slower and safer. If it scales with the algorithmic progress still available, then the more we dig out beforehand, the slower RSI is when it arrives, since RSI is only dangerous when a lot of algorithmic headroom is left. In both cases, slower progress could make RSI faster (as long as compute keeps growing). Note that this is pretty uncertain.
Dropping alignment engineering leaves no account of how models get aligned in practice. Taken at face value, the recommendation implies that well-motivated AI safety people stop working on training models to be aligned, and, if they also leave labs, lose influence there. The argument that prosaic methods are unlikely to remain sufficient through RSI also concedes that they plausibly do. Conditional on that, prosaic methods may be where the marginal gains are largest, since we are not making much visible progress on robustness to RSI.
The field may overstate how little we understand. We know decently well how LLMs work at the level of training dynamics and of what shapes generalisation. What we cannot do is predict their outputs without running them, which is a different claim. E.g., for personas, my guess is roughly this: correlated distributions in pretraining data produce high-level features that are useful for predicting large parts of that data; these features act as generalisation knobs; and instruction fine-tuning biases their activation and reduces how much they condition on context. What more should we aim to understand? The detail of the computations, the robustness of these features to later training?
On the unified-persona assumption (footnote 2), I do not think its explicit version was defended. Split personas should not have surprised anyone, given how much work shows models conditioning their behaviour on contextual cues. Made explicit, the claim would have been that persona training survives later training that optimises directly against it (e.g., hackable environments where hacking is explored and rewarded). With enough training and weak enough regularisation, we should expect the behaviour to change. Persona training is useful when later training does not optimise against it, and for shaping what later training reinforces (e.g., by shaping exploration).
The claim that we do not know how to solve alignment conflates two claims. We know a good deal about what to do, e.g., remove misspecification of the training signal (do not reward hacking, cheating, or deception); reduce reliance on generalisation, by specifying the target behaviour during training or by shaping generalisation (selective generalisation, character training); and get humans to stop training models to be malevolent or selfish. What we do not know is how to do this while staying competitive, and misalignment science helps much less with that second problem.
On implicit work trials, the relevant counterfactual may be capabilities, not misalignment science. Much of the community is not made up of impartial altruists, and the trend looks like it is going the wrong way (less EA over time, less veganism), which is what I would expect from any community growing this fast. Discouraging alignment engineering may not redirect that subset to misalignment science. It may redirect them to capabilities, where the same skills pay better.
I know much less about personas than you do, so I’d be very glad to be proven wrong here, but I feel like there’s a lot that remains to be understood. For example: What counts as one correlated distribution? Is it text produced by one actor (e.g., a single person), a cluster of similar actors (e.g., past AI models or the rationalist community), or something else? How entangled are those distributions with each other? If we post-train the LLM to take on a persona that isn’t represented in the pretraining data, is it going to stitch together features of many different personas, simulate the most similar persona that was represented in pretraining data and learn features that override the behavior of that persona in appropriate ways, or something else? Why does misalignment conditionalize but capabilities don’t? What are the most important factors that determine whether an aligned persona persists or doesn’t persist through conflicting RL pressures? A deep science in the sense Richard Ngo uses the word would probably attempt to tackle more fundamental questions than the ones above, but even the shallower science of personas doesn’t look nearly complete to me.
I think it’s true that a lot of past work foreshadowed split personas, but also, I think it was completely reasonable for people to assume after the Natural Emergent Misalignment paper came out that reward hacking propensities would continue to generalize into a globally misaligned persona instead of conditionalizing into a split-brained grader sycophant. We got evidence that there’s often no emergent misalignment under different training setups a few months later, but a robust theory of personas would have allowed us to predict this right when Anthropic’s paper came out, and I don’t think we had that theory.
Thanks for this post! As a newcomer to AI safety, and seeing some of these debates, reading this helps clarify my thinking of what kind of work I should be doing.
I’m curious where you would place mech interp in this ontology? Two of its major applications seem to be:
Building a deeper understanding of what the models are doing and how they work
Designing more robust evaluations which can catch misalignment that blackbox methods would have missed.
So it seems much more on the ‘Science’ side, but I am hesitant to label all of it as Science? I guess I could imagine a point where mech interp becomes a dominant alignment eval method, such that improving mech interp methods just leads to greater confidence the models are aligned, unlocking continued acceleration of capbalities? Clearly we are not in this world yet, but could be there in 1 year. Still if these methods were good enough that this is actually earned (rather than false) confidence in the models alignment, it might not be entirely unjustified.
Anyway since mech interp seems more on the ‘Science’ side by default, perhaps its also a place where people who have an inclination towards ‘Engineering’ style hill climbing / metric improvement could direct their skills? Engineering performing Mech Interp methods being useful for Science
Thanks for writing this up! I found this post and our own conversations around it very clarifying for my own thoughts.
Like many people I’ve recently become more convinced that a pause is both more urgent and more possible, and so I think the “Building a public evidence base” point is really important. To expand on this, based on recent conversations I’ve had with other researchers, I think a combination of trying to pin the labs to ambitious safety claims / predictions (e.g., “Based on our training and alignment methods, we think our next model will be aligned before we’ve even trained/tested it.”) and then red-teaming this might be really effective. This would then give policy/gov people more reliable signals that alignment is hard and we should pause. (Contrast “X group think it will be safe, but Y group think it will be dangerous” with “X group said it would be safe because of Z, but Y group demonstrated Z is not true”.)
The red-teaming could come in many forms such as: “you claim method X works, but we broke it”, “you’re confident in method X generalising because it has property Y, but we show this isn’t true”, “you’re implicitly relying on some trend in motivational tendencies, but when we actually checked the trendlines they looked bad”, etc. I mean this to say it’s not just red-teaming in the narrow break-the-thing way, but rather red-teaming the claims by doing the relevant science.
I also think one could try doing a sort of hill-climbing on policy/gov people’s reactions to this sort of work (with care of course, you wouldn’t want to be too adversarial / goodhearty). Essentially you do some science to show alignment is harder than expected, show it to the policy makers, if they’re not convinced find out why, then do some more science targeted at that. Obviously it won’t be this simple in practice for a variety of reasons, but I think if people’s theory of change is “build public evidence”, they should try and find out the ways in which their evidence might be failing to resonate with people.
Thanks! IMO this is an even better explanation than the two posts which @Richard_Ngo managed to write...
Note that I consider it to be advocating a pretty different thing than what I’m advocating for, as per this comment.
To the extent that AGI alignment is difficult[1], I think “misalignment science” is a poor term for a field of inquiry that aims to get at the meat of what technical alignment should aim to get at. This is like how it’d be silly to frame the study of how to make lasers as “failure-for-a-random-clump-of-stuff-to-be-a-laser science”. Like, if we want to make lasers, we should ask for [fundamental understanding of light and of materials and understanding of certain specific phenomena and design and fabrication ideas and methods and protocols and specifications of component designs etc] relevant to making lasers — that is, for fields of optics and quantum theory and condensed matter physics and laser science and engineering etc — whereas it makes much less sense to ask for [an understanding of the myriad ways in which a collection of fundamental particles might fail to make up a laser]. Ditto for lenses and particle accelerators and microchips and technologies generally.
And I do think AGI alignment is probably extremely difficult, and so by modus ponens, I think “misalignment science” is a poor term. But I sorta don’t want to defend this part in the present comment. I mostly just want to point out that [using the term “misalignment science” to designate a field of inquiry that captures a lot of what alignment ought to be doing] would contribute to framing the AGI situation as one where there is “normal good/benign behavior” and one studies pathological deviations away from it (for example, medicine has historically been such a field[2]).
That said, I think that your three-point characterization of “misalignment science” is not actually making this imo-mistake. I agree understanding is good and important in AI alignment (though purely technically, ie ignoring governance implications, I’m also very pro alignment engineering work[3]). I agree that security mindset is good and important in AI alignment, and that people would do well to think more about ways their plans might fail[4]. I agree alignment research success should not need to look like having an immediate very practical use case.
But this doesn’t make “misalignment science” a fine term. Furthermore, looking at the sub-approaches you list after your characterization, I think you probably are also making an error in content in this direction indicated by the imo-error in your choice of term.[5]
ie to the extent that it is difficult to have a good transition to a world with artificial generally intelligent systems, or that it is difficult to make AI systems that it is good/fine to hand the future to, or that it is difficult to make AI systems that extend a human’s agency or just benignly do tasks or provide services to humanity, or whatever
I say “historically” because I expect this would change fairly soon in a humane future, with people getting more into longevity/”healthmaxxing”.
I’d only give sth like that the main reason we won’t be disempowered is the present stack of fairly shallow ML research (though tbh I’d include basically everything you list under “misalignment science” under “shallow ML research” as well, but I’d consider alignment engineering to get most of the Shapley here — anyway there are probably disagreements here that are outside the intended scope of the present comment), but I’d also only give sth like that we won’t be disempowered mostly due to other things. (In other words, I’m saying: like we will be disempowered, but ML slop gets credit for a bunch of the remaining .)
Though I want to note that in practice, in the current AI safety community, this will largely look like looking for causes of misalignment exclusively or even just mainly in training pressures, which imo entails already having fundamentally misunderstood the nature of AGI risk. I believe this is a long-standing disagreement that I won’t do justice to in the present comment, but I say some relevant things in this comment.
Btw I feel like your explicit three-point characterization of “misalignment science” and your list of sub-approaches+examples are quite incongruent, at least if I am to take the latter as aiming to span the space decently well. This makes me mostly suspect that your explicit characterization isn’t getting at the field you really have in mind.
I do find it interesting that you stopped at that point and that you didn’t mention any of the agent foundations stuff as to me you also just pitched specific types of agent foundations. Other than that, good frame, I completely agree.
This is at least in part because I’m very poorly read wrt agent foundations, and wouldn’t want to make claims about a field outside my knowledge. But in principle I support this if it builds understanding
Agree with many points but—
the examples given in “Misalignment Science” are… half of it is also mostly engineering? Or: there isn’t that strict boundary between science and engineering, and the science of the type “we came up with a new method to measure the quality of paint applied to widgets” is not of the depth needed
- I dislike the entanglement between Science/Engineering and Alignment/Misalignment. One of the large failures of mainstream AI safety community is the failure to notice when AIs are surprisingly aligned and ask “Why?”. “Alignment faking” paper is the most famous example of this failure.
Thank you. The question of whether AI safety research is net-positive or net-negative was a point of contention that contributed to PauseAI (global) disassociating from PauseAI US, so this has implications beyond personal career choices and the allocation of automated research.
I’ve always had an inkling that the “net-negative” arguments proved too much, and I think this post elucidates the critical distinctions which demonstrate that the “net-negative” argument is true in some cases but not in others.
On the object level, for the type of work that is most likely to help versus hurt: I propose working on figuring out what would help more than hurt.
Almost nobody is funded to figure out what work would solve alignment
Everyone does some as a volunteer. At least doing that is unlikely to accelerate progress, although there are always strange side effects in sharing insights.
Useful distinction, and I agree with a lot of this. A few complications, though.
As others have said here, better understanding of AI behavior/cognition will also help with current alignment goals, so the science speeds things up too. Knowing why a method works is the best route to a better method.
I think it’s very hard to tell which techniques will still help when we hit weak superintelligence, let alone strong. Current alignment techniques may very well keep helping if LLMs get us at least to weak superintelligence. I doubt they’re sufficient, but they may help. So I don’t think the engineering vs. science line tells us which work will turn out to matter.
The capability/alignment dilemma I’ve been most concerned with is on a different axis. I think that Human-like metacognitive skills will reduce LLM slop and aid alignment and capabilities. I’ve been wondering whether the capabilities gain might be worth it.
We are probably going to rely heavily on AI to help with alignment. We’ll try to use it for conceptual questions whether or not the answers are likely to be slop. If we could reduce slop without aiding general reasoning, it would pretty clearly be beneficial. Better metacognition is how humans reduce our slop, but it also helps our capabilities a lot, because knowing when you’re not sure lets you work harder on that part of the problem.
I don’t think there’s a way to avoid dilemmas like this. Capabilities are going to keep improving with or without the help of the risk-concerned. Judging whether we’re helping the odds of alignment more than we’re hurting with acceleration is going to remain hard. Each case should probably be publicly discussed in detail. It’s easy to let motivated reasoning convince you that your favorite alignment approach will be more good than harm. But doing nothing to keep our hands clean is not likely to be the best move, either.
In a somewhat biased defense of alignment engineering: if RSI is near, we will soon have to mostly defer to automated researchers for safety work (alignment engineering, misalignment science, interp, control, evals etc).
It seems likely to me that today’s alignment engineering actually moves the models closer to really trying their best to serve our interests, and that more of this can increase the chance of us actually having aligned automated alignment researchers.
If RSI is near, it seems likely that first automated researchers are very similar to current models, and also that we are only a few generations away, so prosaic alignment work that remains valuable on timescale of ~1 year seems very useful now.
As for worries that prosaic alignment solves issues that would have been warning shots without solving later issues that can actually lead to takeover/extinction, well, after recent events I now think that the kinds of prosaic alignment work we need to do to make today’s models behave better and avoid warning shots will already be pretty close to kind of work needed to also make them behave properly in early automated safety research.
On a similar note, the smarter the smartest trusted model is, more we can rely on control in those early days of RSI.
If RSI is far (which I guess would also imply that models have more time to learn to build untraceable inner thinking / ability to scheme and move away from monitorable CoT entirely, maybe even move away from transformer architecture / even from deep learning), then yes we should pursue long term “solve it for real” works in ambitious mech interp, agent foundations etc.
Good post. I think you are missing two points:
Conditional on “Advancing prosaic alignment advances capabilities”, I think a good reason for someone not to work on alignment engineering is that helping companies that they find evil is bad from a deontological standpoint.
I think you should have mentioned that misalignment science might very well have effects that advance prosaic alignment and therefore advance capabilities. The dual use here seems underemphasised.
As someone with a career in technical AI safety, I think communicating about the threats (to politicians and media) is more impactful/important than technical AI safety research.[1]
AI safety is not (and shouldn’t be) just about technical research. I urge people to actually consider comms. People with or without experience can have impact via meeting with their representatives, see How a cold email got the Finnish government to respond on superintelligence regulation.
I was motivated to write up my experience doing AI safety comms as a technical person because of posts like this one that discuss prioritization (when it’s difficult to have positive impact in technical AI safety research) and use an appendix/footnote to scope out discussion of comms. If it’s difficult to have positive impact in technical AI safety, speaking out / doing advocacy is often a much better way to have impact.
It seems to me that we definitely need more time and are not on track to solve alignment or prevent extinction through technical means even with more investment. I imagine a lot of AI safety researchers also think we need more time. Please be transparent in your communications.
Nice post! It does seem pretty important to explicitly demarcate the specific difference(s) between alignment engineering and misalignment science, for a few reasons:
If we want to incentivize safety work at larger scales, we’ll need more legible criteria for what research is actually helping vs safetywashing.
As AI labor goes into safety work, it probably differentially helps a lot more with the engineering than the science even under the rosiest assumptions, and we can’t mitigate/account for this if we don’t know how to track this.
Another thought: a crude way to describe our situation is by
where:
The main defect in this formula is that doesn’t necessarily tell us anything about how aligned systems will be when they’re first capable of taking over, let alone at strong superintelligence. Maybe we could fix this by instead doing:
where these variables now denote the real vs surface alignment at that critical future point. This is still a crude model, the trajectory over time definitely matters, but maybe there is a nice way to write down these tradeoffs.
Worth considering that not explicitly doing anything AI Safety related could be a good option, given the risk of inadvertently making things worse.
This does not sound like a good scientific practice to me. Understanding why methods work and
understandingrationalizing why they should not are different activities and a goal of “show that X is bad” seems to be the latter. Sometimes alignment approaches really do generalize better than we expected(CoT or even linear probes).Thanks to the “alignment engineering,” we have at least some understanding of what’s going on and partially know why. Alignment is a moving target, and engineers mostly register events, sometimes explaining them in retrospect; that’s why this cannot be called science in principle. The AGI companies invested everything in this “moving target”, putting everything on the line, so we should not expect any other results from them. The “misalignment science” can suffer from the same problems.
In my opinion, the solution isn’t to run after the target, but to rebuild the target itself. The transformer is an extremely simple architecture with just a few inductive biases. We can’t expect aligned behavior from this technology. This was a theoretical idea, but now it’s backed by experience, and experience builds up.
Alignment science is just computer science. If this isn’t possible, then alignment isn’t possible in principle.
I very much agree on the engineering vs science distinction here. Safety is far from an objective concept (though some components of it, like extinction, are) so the thought that it is engineerable seems strange to me, especially when dealing with AI.
There’s definitely a need for a new paradigm, one that acknowledges that AI alignment is not “solvable” in the same way that safety for humans is not solvable.