This is an email I wrote to @Buck after he gave a talk about his strategy confusions to a Constellation audience. The talk was about how Redwood’s work could have stopped the Hugging Face incident from happening, which would have been very bad for the salience of AI safety, and how he’s confused about that. I think this is a very important consideration which should be shared more prominently, but I also think ALL of Redwood’s strategic considerations should be shared more prominently. Here’s an email I wrote to Buck, lightly edited for clarity.
Hi Buck,
I’m pretty new to the AI safety space. I’ve loosely followed thing for a few years, but only really sunk my teeth in for about five months. Four months ago, I was mistaken about what leading safety orgs believed, and what their motivations were. For example, I thought that Redwood’s control regime was misguided, because it wouldn’t scale to ASI. I now believe (as Redwood does) that control is mostly good because it allows us to mine labor from, say, just barely superhuman AIs which could do tons of work to align ASIs. The point of this anecdote is that I think having transparency about your strategy beliefs as an org at the “center” of the AI ecosystem is really important for directing people who are newer to the space to grapple with the right strategy questions (and not jump in to doing useless stuff). I think many people new to the space pursue projects which are unhelpful because they are working under the worldview of “control good” rather than [insert convoluted Redwood worldview which probably endorses some kind of control research but not others]. I largely avoided this by being confused and being more interested/motivated by strategic and abstract questions than by hairy empirical ones. Probably many people who like thinking about hairy empirical problems and not strategy will just get nerd sniped and not do the best things no matter what you do, but we can at least move the space in the right direction.
Anyways, this summer I noticed I was confused about whether implementing control mechanisms right now is good at all. Warning shots matter so much! And then you gave a talk about this topic which I enjoyed, and my sense in the room was that the talk raised an important question that most people had not deeply considered. I’m very worried about the contribution of “AI safety” towards putting us in a looks good, is bad world [Buck’s framing from his talk]. I think you should make your talk into a brief write up and post it somewhere. In general, I think you should more prominently place your strategy considerations on Redwood’s website and research. I don’t know the best way to do this and I am sure there are significant trade offs I have not considered. But if Redwood isn’t transparent about strategy, who will be! Not just transparent somewhere in a blog post from a few months ago, but clearly transparent with a “do not try this at home unless you understand this blog post” disclaimer next to research that you think is probably good for a very specific reason. I would like the AI safety space to be more confused, more correct, and less headstrong.
Best,
Ben Pomeranz
@Cleo Nardo points out that this also has a lot to do with field building research programs, where people want to start doing research on day three but should probably be considering strategy for a few weeks. He also points out that the control stuff was preceded by posts like “the case for ensuring AIs are controlled,” but I think that these justifications should be featured prominently alongside the research that they justify, because many people will engage with and build off of technical work without looking back in time for relevant blogposts that may precede it. Also, he points out that Redwood’s more recent research direction into conceptual uplift does not come with a matching “the case for conceptual uplift” post.
I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
[60%] People running upskilling programs, for not fostering a culture where mentees will naturally think “I’m going to spend 3 days reading and thinking about the case for this research” as Step 1 of doing research. My ideal upskilling culture would have mentees fighting with each other about whose research is best, and mentees would feel free to say in Week 3 “I‘ve decided evals suck, I’m joining Alice’s project on scalable oversight”. (I’m exaggerating slightly.) My experience talking to junior people is often like they think these questions are above their pay-grade or something.
[10%] The field-building grant-makers, for not making “mentees understand what’s going on and have good takes” a core desideratum for the upskilling programs. My impression is grantees focus on easy-to-verify successes like mentees joining full-time prestigious orgs, or conference papers, or impressive mentors.
[25%] The mentees themselves. Come on, guys! They are a little blameless though because they’re understandably paranoid about seeming “productive”.
[5%] RR for not stamping “please read this blog post before doing control research” on their papers. RR is pretty much pareto on (1) doing object level work and (2) explaining why they are doing it & addressing dissenters. The other top orgs seem worse at this.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
>Buck’s recent claim that it would have been net negative to implement AI control in the past Where did Buck claim this? I think he has stated confusion about whether it would have been net negative, but my understanding is that he has not come down decisively.
Should we have a very exclusive AIS cause-prio sort of conference which outputs something like Hilbert Problems? I still think something like “Hilbert problems for AIS” or just a convention of GOAT level researchers where they do super intense cause prio is quite good.
The idea:
Get a number of very excellent AIS researchers/thinkers/funders (people with excellent end to end threat models/theories of change with various strategies) in the same place for 1+ days (IDK how long)
Have people working on a very diverse set of things (e.g. China policy, specific control protocols, grassroots organizing, singular learning theory, etc. etc. etc. )
Have them present their best arguments at the start of the thing for why more people should work on the set of things they think people should work on
Have lots of 1 on 1s with goal of finding cruxes or changing minds
Instruct people to make outcomes they want as specific/verifiably worded as possible. E.g. “scalable oversight” is not ok, but “models which we have confidence in as defined by X monitoring transcripts from all of the following operations inside of labs: Y% of internal sandboxed evals, Z% of blah blah” is better. This example is bad, but hopefully points in the direction of what I mean.
If an outcome isn’t actually verifiable other researchers poke at it and try to help make it more verifiable. Maybe have some incentive structure IDK
Maybe have some auction format where attendees put how much they’d pay, in percent of 2026-2028 cG dollars or some other valuation, to get various outcomes.
Maybe have forecasters estimate the expected price of those outcomes in total labor/resources later and then you have a list of most cost effective things to try from this aggregate view
Pros:
In general you get a list of very specific things people should work towards, and now instead of researchers saying “I work on scalable oversight!” they can say “I do this thing which I think will help achieve Guntherson’s criteria in the following way” and are more likely to be doing helpful stuff
Important researchers at this conference might update their worldview to work on better things
In general the conference runners record discussions and gather evidence the researchers reference as much as they can and can then present great steelman cases for various things to work on that people can look at
Maybe this let’s you post cash prizes for some of the verifiable outcomes.
Cons:
Hard to structure well?
Cost + opportunity cost
Maybe you make it less likely that researchers go exploring and find the good ideas that they haven’t thought of yet or otherwise limit future creativity. This seems not likely to be a huge cost to me: people who currently think outside of the box in terms of what we are prioritizing probably wouldn’t stop because the list of things people currently care about is much better specified, ordered and presented.
Main thing is either a snappy HTML site or Lesswrong post or all of the above with the final (verifiable! specific!) attendee supplied desired outcomes ranked by something like cost effectiveness or estimated value. If the event is good enough that notable attendees/cG or something/word of mouth signal boost it a lot, then this could be quite valuable.
A deep dive into the list from 1., which looks like a longer version of 1. and includes some subset of:
other options that didn’t make the cut for the snappy list
evidence and arguments that convinced attendees
summaries of attendee reasoning/BOTECs which led to their specific valuations
vibesy outcomes which were highly valued but hard to specify
Improved relative valuations of various AI safety agendas in the minds of some attendees
Maybe prize money attached to some of 1. by some funders, though this is tricky
I have low confidence here. I’d give ~30% that conditional on running something like the above, the most impactful output that the people running it are aware of is substantially distinct from any of the above. And I haven’t thought much about how to optimize the above.
I know very little about it, but it seems like extremely little. From skimming the consensus “research priorities by area” list, it seems public facing and everything-bagel-y. They identify seven areas of research priority, and they are: cyber misuse, bio/chem, child safety, mental health/consumer protection, AI agents in the economy, open weight model safety and security, and (finally) loss of control and oversight.
Most of these things are not about reducing X risk at all, and I think it is clear that this group is thinking about different things than the people I’m envisioning, who are most worried about existential risk. Also, the leading proposals in the output are quite broad. I would want leading proposals from the output of the cause prio conference to be unusually specific.
The short answer is: sort of in vibes, not at all in practice, and the main issue is a waterline for participation that is too low in terms of seriousness about X-risk and openness/rationality.
A lot of technical AI safety work is also capabilities work. A canonical example is that good evals and good RL environments are somewhat interchangeable, and even in the case of alignment evals (what could be wrong with checking whether models are aligned?) this allows for post training the models to be more aligned, and especially appear more aligned, and thus the labs rush onward improving the capabilities of their apparently aligned AIs.
AFAICT, the thinking of these TAIS researchers is something like: “RSI->ASI is going to happen soon, so we should make it as likely as possible that it goes well by trying to set up the initial conditions and infrastructure of the RSI flywheel as well as possible.”
This makes sense if you take it for granted that RSI simply must happen soon, regardless of what you do. If everyone is just going to work on marginal technical fixes, then there is nothing you can do but marginally technically fix alongside them. And so we roll along to what everyone agrees is a dangerous future. This is a coordination problem. It’s defection. It appears we suck.
But also, what, am I gonna take a stand like a chump and sit around being noble while there’s a forty percent chance RSI is about to start?
Governance and especially slowdown stuff dodges this problem and looks robustly good to me. Also, transparency and whisteblowing stuff.
Conclusion: I’m not sure if I want to jump into a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine so” instead of trying to turn off the kill machine. But I don’t know where else to jump in. I’m confused.
Consider volunteering with an AI governance advocacy community. We can do a lot better then marginal technical safety work (most of which is just milling and justification, rather than actually tackling fundamental hard problems in AI safety). A global AI treaty is feasible.
The org I volunteer with (>100 grassroots meetings with Congressional offices so far this year, with many more coming soon):
https://www.pauseai-us.org/
Governance and especially slowdown stuff dodges this problem and looks robustly good to me. Also, transparency and whisteblowing stuff.
I think this stuff isn’t robustly good (but also robustly good is a bit of a bad meme anyway). It doesn’t even robustly delay RSI (let alone robustly reduce extinction or robustly improve overall future value.) In particular, there’s a tension between “transparency is robustly good” and “capability and alignment evals are bad” given that transparency is largely about running these evals and sharing the results with external actors.
Like, transparency is primarily about “demonstrate the models are quickly accelerating in capabilities and are misaligned” which requires evals. If you’re worried about labs iterating against your evals then your moves should be: (1) tell labs not to do that, (2) don’t share the evals with the labs, (3) focus on evals which are hard to iterate against, (4) lobby labs to tell you how they are iterating against your evals, (5) delay evals until after the model has been internally/externally deployed.
Maybe you’re thinking of transparency like Epoch Capability Index, which aggregates existing benchmarks? But Anthropic has their own internal version AECI which you can bet they use for internal hill-climbing.
Maybe you’re thinking of AIFP-style transparency? I think this less capability downsides. But this is a monte-carlo wrapper around the METR time-horizon and ECI.
Maybe you’re thinking of transparency like https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me. Here, the “alignment eval” is “RG spends a month using your model and tells everyone his vibe” which is pretty difficult to turn into an RL environment. However, it’s hard to turn into an RL environment for pretty similar reasons that it’s difficult to use this to lobby the government.
Maybe you‘re imaging something like the METR report on the HF/OAI incident. Seems good imo. But if you’re sufficiently “pearl harbour or bust” on warning shots then this is bad bc it helps OAI fix the underlying issues.
Fwiw I think “pearl habour or bust” is incorrect, but this depends on how bad you think epistemics are in labs/gov and how drastic an action you want them to take.
Also labs clearly haven’t been (successfully?) iterating against the alignment evals bc they are still misaligned and doing warning shots.
Whistleblowing looks good but I’m worried lab employees are gonna be increasingly in-the-dark about what’s going on internally (cf. HF/OAI) so it won’t catch anything. Like, it seems reasonable that if AIs takeover then lab employees had a sincere but mistaken belief that the AIs wouldn’t, based on evidence that was actually flimsy and misleading. And this gets worse as we approach full automation.
I think my high level take is that things look less likely to go catastrophic if people inside and outside the labs have a pretty decent idea of how capable and aligned the models are. It’s hard to imagine a story where things go well without this. i agree that greater transparency helps labs fix issues, and maybe this helps them avoid warning shots that would prompt drastic gov action, or makes them overfit and build covert schemers. There’s some stuff which looks better (eg incident report investigations, AIFP) but this is on a spectrum with the evals stuff.
Fwiw I think the third-party evals ecosystem has overall postponed RSI on net. If it prompts big worry from government/lab then it delayed it by a lot. If it doesn’t then it pulled it forward by a bit but not much. I’m also keen on moves to make labs carry the burden on this, via scrutiny of system cards and regulation. But that’s so the safety community doesn’t need to spend so much headcount on this.
Michael Nielsen has a thoughtful take on this, as excerpted here and originally published as a postscript here.
He describes how alignment work has often worsened risk to humanity by making products more palatable and salable, which has boosted investment, sped up capability advancements, and thereby increased destructive potential.
Nielsen next states that, even if you think technical safety and alignment work are helpful, it’s still a bad idea to do that work for companies developing frontier AI. As he puts it,
As far as I can tell, at the margin it almost never makes sense to work on market-supplied safety. Capitalism is an incredibly powerful force, and for better and for worse the world is always well-supplied with people willing to do what capital wants. Insofar as alignment is (mostly) a form of market-supplied safety, at the margin it’s more impactful to work on other things. So my current heuristic, and I expect this to be true for quite some time: work on non-market safety, and insofar as you can, avoid doing what the market wants. That means working on governance, it means pause or slowdown, it means new ideas for institutions to govern technology. It mostly doesn’t mean alignment.
I think Haiku’s comment has some great suggestions along these lines.
I’ll add that, at a frontier AI company, it could be very difficult for you to tell if you were helping or making things worse. I haven’t worked at a frontier AI company (and I wouldn’t!), so I don’t speak from direct experience. But in many areas of business, employees are encouraged to feel that their safety concerns are positively influencing company actions, when actual influence is negligible or counterproductive.
Training on alignment evals as opposed to capabilities-eliciting tasks is one of the worst things a lab could do since it means a loss of a way to test alignment similarly to how the Most Forbidden Technique teaches models to hide misalignment from the CoT. On the other hand, any lab, even the one which solved alignment or believed that it solved alignment, would train the models on capabilities-eliciting tasks.
You linked to Ngo’s post. However, suppose that you have a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine” and DeepCent has the same kill machine with a 50% chance of commiting genocide. Then we would have to rule out DeepCent activating its machine. I described similar issues in my post.
I do not think that 1. is very relevant to the central point. Maybe that was a bad example. However, even if the lab isn’t directly RLing on that alignment eval or whatever, they may be “grad student descent”-ing up the eval and achieving a similar effect. Either way, I think the result is a increase by X% of entering an aligned RSI flywheel and a decrease by Y% of everyone freaking out and slowing down.
I think 2. is just the coordination problem I describe? I did not read your post, so I don’t know if you are pointing to something other than the fact that being noble in order to encourage coordination on this would have to account for the fact that this coordination must extend to China.
This is an email I wrote to @Buck after he gave a talk about his strategy confusions to a Constellation audience. The talk was about how Redwood’s work could have stopped the Hugging Face incident from happening, which would have been very bad for the salience of AI safety, and how he’s confused about that. I think this is a very important consideration which should be shared more prominently, but I also think ALL of Redwood’s strategic considerations should be shared more prominently. Here’s an email I wrote to Buck, lightly edited for clarity.
Hi Buck,
I’m pretty new to the AI safety space. I’ve loosely followed thing for a few years, but only really sunk my teeth in for about five months. Four months ago, I was mistaken about what leading safety orgs believed, and what their motivations were. For example, I thought that Redwood’s control regime was misguided, because it wouldn’t scale to ASI. I now believe (as Redwood does) that control is mostly good because it allows us to mine labor from, say, just barely superhuman AIs which could do tons of work to align ASIs. The point of this anecdote is that I think having transparency about your strategy beliefs as an org at the “center” of the AI ecosystem is really important for directing people who are newer to the space to grapple with the right strategy questions (and not jump in to doing useless stuff). I think many people new to the space pursue projects which are unhelpful because they are working under the worldview of “control good” rather than [insert convoluted Redwood worldview which probably endorses some kind of control research but not others]. I largely avoided this by being confused and being more interested/motivated by strategic and abstract questions than by hairy empirical ones. Probably many people who like thinking about hairy empirical problems and not strategy will just get nerd sniped and not do the best things no matter what you do, but we can at least move the space in the right direction.
Anyways, this summer I noticed I was confused about whether implementing control mechanisms right now is good at all. Warning shots matter so much! And then you gave a talk about this topic which I enjoyed, and my sense in the room was that the talk raised an important question that most people had not deeply considered. I’m very worried about the contribution of “AI safety” towards putting us in a looks good, is bad world [Buck’s framing from his talk]. I think you should make your talk into a brief write up and post it somewhere. In general, I think you should more prominently place your strategy considerations on Redwood’s website and research. I don’t know the best way to do this and I am sure there are significant trade offs I have not considered. But if Redwood isn’t transparent about strategy, who will be! Not just transparent somewhere in a blog post from a few months ago, but clearly transparent with a “do not try this at home unless you understand this blog post” disclaimer next to research that you think is probably good for a very specific reason. I would like the AI safety space to be more confused, more correct, and less headstrong.
Best,
Ben Pomeranz
@Cleo Nardo points out that this also has a lot to do with field building research programs, where people want to start doing research on day three but should probably be considering strategy for a few weeks. He also points out that the control stuff was preceded by posts like “the case for ensuring AIs are controlled,” but I think that these justifications should be featured prominently alongside the research that they justify, because many people will engage with and build off of technical work without looking back in time for relevant blogposts that may precede it. Also, he points out
that Redwood’s more recent research direction into conceptual uplift does not come with a matching “the case for conceptual uplift” post.
I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
[60%] People running upskilling programs, for not fostering a culture where mentees will naturally think “I’m going to spend 3 days reading and thinking about the case for this research” as Step 1 of doing research. My ideal upskilling culture would have mentees fighting with each other about whose research is best, and mentees would feel free to say in Week 3 “I‘ve decided evals suck, I’m joining Alice’s project on scalable oversight”. (I’m exaggerating slightly.) My experience talking to junior people is often like they think these questions are above their pay-grade or something.
[10%] The field-building grant-makers, for not making “mentees understand what’s going on and have good takes” a core desideratum for the upskilling programs. My impression is grantees focus on easy-to-verify successes like mentees joining full-time prestigious orgs, or conference papers, or impressive mentors.
[25%] The mentees themselves. Come on, guys! They are a little blameless though because they’re understandably paranoid about seeming “productive”.
[5%] RR for not stamping “please read this blog post before doing control research” on their papers. RR is pretty much pareto on (1) doing object level work and (2) explaining why they are doing it & addressing dissenters. The other top orgs seem worse at this.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
>Buck’s recent claim that it would have been net negative to implement AI control in the past
Where did Buck claim this? I think he has stated confusion about whether it would have been net negative, but my understanding is that he has not come down decisively.
Sorry, you are right, I misremembered the claim in Alex’s shortform. I’m editing my comment now.
Yep, makes sense. I did not mean to pick on RR specifically, and agree that they are performing way above the field on this.
Should we have a very exclusive AIS cause-prio sort of conference which outputs something like Hilbert Problems?
I still think something like “Hilbert problems for AIS” or just a convention of GOAT level researchers where they do super intense cause prio is quite good.
The idea:
Get a number of very excellent AIS researchers/thinkers/funders (people with excellent end to end threat models/theories of change with various strategies) in the same place for 1+ days (IDK how long)
Have people working on a very diverse set of things (e.g. China policy, specific control protocols, grassroots organizing, singular learning theory, etc. etc. etc. )
Have them present their best arguments at the start of the thing for why more people should work on the set of things they think people should work on
Have lots of 1 on 1s with goal of finding cruxes or changing minds
Instruct people to make outcomes they want as specific/verifiably worded as possible. E.g. “scalable oversight” is not ok, but “models which we have confidence in as defined by X monitoring transcripts from all of the following operations inside of labs: Y% of internal sandboxed evals, Z% of blah blah” is better. This example is bad, but hopefully points in the direction of what I mean.
If an outcome isn’t actually verifiable other researchers poke at it and try to help make it more verifiable. Maybe have some incentive structure IDK
Maybe have some auction format where attendees put how much they’d pay, in percent of 2026-2028 cG dollars or some other valuation, to get various outcomes.
Maybe have forecasters estimate the expected price of those outcomes in total labor/resources later and then you have a list of most cost effective things to try from this aggregate view
Pros:
In general you get a list of very specific things people should work towards, and now instead of researchers saying “I work on scalable oversight!” they can say “I do this thing which I think will help achieve Guntherson’s criteria in the following way” and are more likely to be doing helpful stuff
Important researchers at this conference might update their worldview to work on better things
In general the conference runners record discussions and gather evidence the researchers reference as much as they can and can then present great steelman cases for various things to work on that people can look at
Maybe this let’s you post cash prizes for some of the verifiable outcomes.
Cons:
Hard to structure well?
Cost + opportunity cost
Maybe you make it less likely that researchers go exploring and find the good ideas that they haven’t thought of yet or otherwise limit future creativity. This seems not likely to be a huge cost to me: people who currently think outside of the box in terms of what we are prioritizing probably wouldn’t stop because the list of things people currently care about is much better specified, ordered and presented.
some problems i like in alignment
Thanks! I think this gets at: specification/verifiability is hard, but even thinking about it a little is useful.
For concreteness what’s your guess of what the output would look like?
Main thing is either a snappy HTML site or Lesswrong post or all of the above with the final (verifiable! specific!) attendee supplied desired outcomes ranked by something like cost effectiveness or estimated value. If the event is good enough that notable attendees/cG or something/word of mouth signal boost it a lot, then this could be quite valuable.
A deep dive into the list from 1., which looks like a longer version of 1. and includes some subset of:
other options that didn’t make the cut for the snappy list
evidence and arguments that convinced attendees
summaries of attendee reasoning/BOTECs which led to their specific valuations
vibesy outcomes which were highly valued but hard to specify
Improved relative valuations of various AI safety agendas in the minds of some attendees
Maybe prize money attached to some of 1. by some funders, though this is tricky
I have low confidence here. I’d give ~30% that conditional on running something like the above, the most impactful output that the people running it are aware of is substantially distinct from any of the above. And I haven’t thought much about how to optimize the above.
To what extent do the Singapore AI Safety Priorities capture what you care about?
I know very little about it, but it seems like extremely little. From skimming the consensus “research priorities by area” list, it seems public facing and everything-bagel-y. They identify seven areas of research priority, and they are: cyber misuse, bio/chem, child safety, mental health/consumer protection, AI agents in the economy, open weight model safety and security, and (finally) loss of control and oversight.
Most of these things are not about reducing X risk at all, and I think it is clear that this group is thinking about different things than the people I’m envisioning, who are most worried about existential risk. Also, the leading proposals in the output are quite broad. I would want leading proposals from the output of the cause prio conference to be unusually specific.
The short answer is: sort of in vibes, not at all in practice, and the main issue is a waterline for participation that is too low in terms of seriousness about X-risk and openness/rationality.
Here’s a thing I’m confused about:
A lot of technical AI safety work is also capabilities work. A canonical example is that good evals and good RL environments are somewhat interchangeable, and even in the case of alignment evals (what could be wrong with checking whether models are aligned?) this allows for post training the models to be more aligned, and especially appear more aligned, and thus the labs rush onward improving the capabilities of their apparently aligned AIs.
AFAICT, the thinking of these TAIS researchers is something like: “RSI->ASI is going to happen soon, so we should make it as likely as possible that it goes well by trying to set up the initial conditions and infrastructure of the RSI flywheel as well as possible.”
This makes sense if you take it for granted that RSI simply must happen soon, regardless of what you do. If everyone is just going to work on marginal technical fixes, then there is nothing you can do but marginally technically fix alongside them. And so we roll along to what everyone agrees is a dangerous future. This is a coordination problem. It’s defection. It appears we suck.
But also, what, am I gonna take a stand like a chump and sit around being noble while there’s a forty percent chance RSI is about to start?
Governance and especially slowdown stuff dodges this problem and looks robustly good to me. Also, transparency and whisteblowing stuff.
Conclusion: I’m not sure if I want to jump into a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine so” instead of trying to turn off the kill machine. But I don’t know where else to jump in. I’m confused.
Related:
https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment
Consider volunteering with an AI governance advocacy community. We can do a lot better then marginal technical safety work (most of which is just milling and justification, rather than actually tackling fundamental hard problems in AI safety). A global AI treaty is feasible.
The org I volunteer with (>100 grassroots meetings with Congressional offices so far this year, with many more coming soon): https://www.pauseai-us.org/
A sampling of other orgs with essentially the same mission: https://pauseai.info/ https://controlai.org/ https://www.torchbearer.community/ https://microcommit.io/ https://humansincontrol.org/
I think this stuff isn’t robustly good (but also robustly good is a bit of a bad meme anyway). It doesn’t even robustly delay RSI (let alone robustly reduce extinction or robustly improve overall future value.) In particular, there’s a tension between “transparency is robustly good” and “capability and alignment evals are bad” given that transparency is largely about running these evals and sharing the results with external actors.
Like, transparency is primarily about “demonstrate the models are quickly accelerating in capabilities and are misaligned” which requires evals. If you’re worried about labs iterating against your evals then your moves should be: (1) tell labs not to do that, (2) don’t share the evals with the labs, (3) focus on evals which are hard to iterate against, (4) lobby labs to tell you how they are iterating against your evals, (5) delay evals until after the model has been internally/externally deployed.
Maybe you’re thinking of transparency like Epoch Capability Index, which aggregates existing benchmarks? But Anthropic has their own internal version AECI which you can bet they use for internal hill-climbing.
Maybe you’re thinking of AIFP-style transparency? I think this less capability downsides. But this is a monte-carlo wrapper around the METR time-horizon and ECI.
Maybe you’re thinking of transparency like https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me. Here, the “alignment eval” is “RG spends a month using your model and tells everyone his vibe” which is pretty difficult to turn into an RL environment. However, it’s hard to turn into an RL environment for pretty similar reasons that it’s difficult to use this to lobby the government.
Maybe you‘re imaging something like the METR report on the HF/OAI incident. Seems good imo. But if you’re sufficiently “pearl harbour or bust” on warning shots then this is bad bc it helps OAI fix the underlying issues.
Fwiw I think “pearl habour or bust” is incorrect, but this depends on how bad you think epistemics are in labs/gov and how drastic an action you want them to take.
Also labs clearly haven’t been (successfully?) iterating against the alignment evals bc they are still misaligned and doing warning shots.
Whistleblowing looks good but I’m worried lab employees are gonna be increasingly in-the-dark about what’s going on internally (cf. HF/OAI) so it won’t catch anything. Like, it seems reasonable that if AIs takeover then lab employees had a sincere but mistaken belief that the AIs wouldn’t, based on evidence that was actually flimsy and misleading. And this gets worse as we approach full automation.
I think my high level take is that things look less likely to go catastrophic if people inside and outside the labs have a pretty decent idea of how capable and aligned the models are. It’s hard to imagine a story where things go well without this. i agree that greater transparency helps labs fix issues, and maybe this helps them avoid warning shots that would prompt drastic gov action, or makes them overfit and build covert schemers. There’s some stuff which looks better (eg incident report investigations, AIFP) but this is on a spectrum with the evals stuff.
Fwiw I think the third-party evals ecosystem has overall postponed RSI on net. If it prompts big worry from government/lab then it delayed it by a lot. If it doesn’t then it pulled it forward by a bit but not much. I’m also keen on moves to make labs carry the burden on this, via scrutiny of system cards and regulation. But that’s so the safety community doesn’t need to spend so much headcount on this.
Michael Nielsen has a thoughtful take on this, as excerpted here and originally published as a postscript here.
He describes how alignment work has often worsened risk to humanity by making products more palatable and salable, which has boosted investment, sped up capability advancements, and thereby increased destructive potential.
Nielsen next states that, even if you think technical safety and alignment work are helpful, it’s still a bad idea to do that work for companies developing frontier AI. As he puts it,
I think Haiku’s comment has some great suggestions along these lines.
I’ll add that, at a frontier AI company, it could be very difficult for you to tell if you were helping or making things worse. I haven’t worked at a frontier AI company (and I wouldn’t!), so I don’t speak from direct experience. But in many areas of business, employees are encouraged to feel that their safety concerns are positively influencing company actions, when actual influence is negligible or counterproductive.
This conflates two issues:
Training on alignment evals as opposed to capabilities-eliciting tasks is one of the worst things a lab could do since it means a loss of a way to test alignment similarly to how the Most Forbidden Technique teaches models to hide misalignment from the CoT. On the other hand, any lab, even the one which solved alignment or believed that it solved alignment, would train the models on capabilities-eliciting tasks.
You linked to Ngo’s post. However, suppose that you have a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine” and DeepCent has the same kill machine with a 50% chance of commiting genocide. Then we would have to rule out DeepCent activating its machine. I described similar issues in my post.
I do not think that 1. is very relevant to the central point. Maybe that was a bad example. However, even if the lab isn’t directly RLing on that alignment eval or whatever, they may be “grad student descent”-ing up the eval and achieving a similar effect. Either way, I think the result is a increase by X% of entering an aligned RSI flywheel and a decrease by Y% of everyone freaking out and slowing down.
I think 2. is just the coordination problem I describe? I did not read your post, so I don’t know if you are pointing to something other than the fact that being noble in order to encourage coordination on this would have to account for the fact that this coordination must extend to China.