To be clear, when writing this post, recruiting any mathematicians to alignment work was much more of a stretch goal.
The low-hanging fruit I’m trying to point to is that there are droves of academics currently considering going to work at frontier AI labs without having even heard the basic arguments for x-risk. I think it would be much much easier to convince a substantial fraction of them to do NOT THAT, than to direct them to do any active technical alignment work.
Having mulled this slightly more… I don’t actually have a very clear idea on what would work here (in the sense of “be memetically successful enough to reach a meaningful number of people.”)
I think there’s a moderately effortful job here of “figure out how to communicate about this in terms the math community will be motivated by and engage with” which can only really be done by someone in the math community.
I’m curious, when you imagine doing “the obvious thing”, what ends up being hard or not working? (presumably there’s a reason you wrote this post targeted at LW users than some other post targeted at math people?)
The “obvious thing” that I have done is state my views on doom to people when it comes up and point them to IABIED, and I think it has largely been successful in causing ~20 people I have personal relationships with to arrive at p(doom) > 5%.
The next “obvious thing” is to scale this up by a factor of 10 to people I’m only vaguely acquainted with, and to write more publicly in a mathematician-facing way as you suggested. These feel both costly, risky, and less effective in ways that the previous thing was not. I am still considering doing it, but need to sit down with my aversions for a while. Some examples: these conversations sometimes require me to be forceful/activist to keep people in my frame instead of sliding off into more comfortable topics, and they plausibly only work because I have significant personal rapport already with a colleague or student.
I would certainly prefer to live in a world where 1 in 20 mathematicians are already awake and I can just create common knowledge that the private conversation thing works, and then not do anything more effortful. As far as I know the only other academic mathematician who writes about Doom is Jacob Tsimerman, who just won the Fields Medal and left to do safety at OpenAI. He would certainly be a person who could do the broadcasting thing (and is doing it already, but in a less forceful way than I’m imagining).
(This is probably not helpful, but just want to note that I’m interested, in a “hobbyist” sense, in the problem of the general kind of confrontation you’re mentioning, under the moniker “confrontation-worthy empathy”; if you happen to want to chat about it I’m interested, e.g. to consider different ideas for making the process more wholesome and similar.)
I would be interested in having this conversation! It seems to me that the kind of confrontation I’m mentioning is a much easier version of the thing you’re pointing to? I tend to just point amateur levels of unconditional positive regard at people and under pre-conditions of existing trust conversations mostly go well. I imagine scaling up the operation requires more finesse, but even so it’s hard to imagine needing the amount of skill you describe there. I’ll message you.
Partly, I am kinda assuming people have a harder time wrapping their brain around “what not to do” vs “what to do”, and I think you’ll be more successful with that goal if you present a positive vision of how to orient to x-risk that speaks their language.
But, yeah that is a pretty different framing that would be suspicious if it didn’t look different at all from where I was pointing.
Partly, I am kinda assuming people have a harder time wrapping their brain around “what not to do” vs “what to do”
IDK, there being problems that are hard to fix but easy to make worse seems like a sort of common sense situation? Though I’m struggling to think of any good examples off the top of my head. Also just “we aren’t ready for the machine god, don’t help them build it faster” is super simple?
I guess math very much isn’t like this. You can generally make at least some progress on a problem, but you can’t really make a problem worse except by giving bad advice to others working on it?
I think there’s something especially bad about AI, where people disagree a lot or are confused about what exactly the problems are, and have a lot of motivated cognition towards thinking “if they’re working on it, it’ll go better instead of worse.”
(i.e. if people were right that the main problem is misuse, going to work at a lab is more reasonable)
There’s something additionally hard about the existential stakes, where it’s hard to admit the problem is real if you don’t feel like you have some way of engaging with it.
E.g. Musk seems like mega-biased toward taking action, with bad results (responsible for like a quarter of frontier AI lab stuff, in some sense). The politician’s syllogism bites real hard when accelerating bad stuff is 1. convergent and 2. much easier than accelerating good stuff.
“Please read IABIED” is my first thought, but sounds like you’ve probably tried that? Maybe there should be a TLDR aimed at mathematicians? We could ask Yudkowsky to write one?
I think recommending a book just doesn’t work without a long preamble or some unusual social leverage. I’ve had better success recommending online articles, my current go-to is Scott’s book review of IABIED, which is still tl;dr and has a bunch of missing context.
The main thing a good first-intro has to do, honestly, is to be as short as possible while explaining instrumental convergence, orthogonality, and countering some representative examples of crazy moon logic that people come up with initially.
Thanks for the tip! On first pass I think it is more measured and comprehensive but less fun-to-read than Scott’s review so I would likely not try that instead.
I’d be curious if you (or anybody else) want to test to seeif my intro is superior[1]. I weakly think it would be, but I’m pretty unconfident and there are good reasons for me to share my guide with people I know over other people’s guides regardless of objective merit.
I deliberately wrote the intro to avoid many mistakes I saw in other intros (eg too many extraneous details, appealing to overly complex analogies, shadowboxing insider objections, or “going meta” in an intro post).
To be clear, when writing this post, recruiting any mathematicians to alignment work was much more of a stretch goal.
The low-hanging fruit I’m trying to point to is that there are droves of academics currently considering going to work at frontier AI labs without having even heard the basic arguments for x-risk. I think it would be much much easier to convince a substantial fraction of them to do NOT THAT, than to direct them to do any active technical alignment work.
Having mulled this slightly more… I don’t actually have a very clear idea on what would work here (in the sense of “be memetically successful enough to reach a meaningful number of people.”)
I think there’s a moderately effortful job here of “figure out how to communicate about this in terms the math community will be motivated by and engage with” which can only really be done by someone in the math community.
I’m curious, when you imagine doing “the obvious thing”, what ends up being hard or not working? (presumably there’s a reason you wrote this post targeted at LW users than some other post targeted at math people?)
The “obvious thing” that I have done is state my views on doom to people when it comes up and point them to IABIED, and I think it has largely been successful in causing ~20 people I have personal relationships with to arrive at p(doom) > 5%.
The next “obvious thing” is to scale this up by a factor of 10 to people I’m only vaguely acquainted with, and to write more publicly in a mathematician-facing way as you suggested. These feel both costly, risky, and less effective in ways that the previous thing was not. I am still considering doing it, but need to sit down with my aversions for a while. Some examples: these conversations sometimes require me to be forceful/activist to keep people in my frame instead of sliding off into more comfortable topics, and they plausibly only work because I have significant personal rapport already with a colleague or student.
I would certainly prefer to live in a world where 1 in 20 mathematicians are already awake and I can just create common knowledge that the private conversation thing works, and then not do anything more effortful. As far as I know the only other academic mathematician who writes about Doom is Jacob Tsimerman, who just won the Fields Medal and left to do safety at OpenAI. He would certainly be a person who could do the broadcasting thing (and is doing it already, but in a less forceful way than I’m imagining).
(This is probably not helpful, but just want to note that I’m interested, in a “hobbyist” sense, in the problem of the general kind of confrontation you’re mentioning, under the moniker “confrontation-worthy empathy”; if you happen to want to chat about it I’m interested, e.g. to consider different ideas for making the process more wholesome and similar.)
I would be interested in having this conversation! It seems to me that the kind of confrontation I’m mentioning is a much easier version of the thing you’re pointing to? I tend to just point amateur levels of unconditional positive regard at people and under pre-conditions of existing trust conversations mostly go well. I imagine scaling up the operation requires more finesse, but even so it’s hard to imagine needing the amount of skill you describe there. I’ll message you.
Yeah, I think it should be much less hard for people not already embedded in working on AI; just seems kinda related.
Yeah fair.
Partly, I am kinda assuming people have a harder time wrapping their brain around “what not to do” vs “what to do”, and I think you’ll be more successful with that goal if you present a positive vision of how to orient to x-risk that speaks their language.
But, yeah that is a pretty different framing that would be suspicious if it didn’t look different at all from where I was pointing.
IDK, there being problems that are hard to fix but easy to make worse seems like a sort of common sense situation? Though I’m struggling to think of any good examples off the top of my head. Also just “we aren’t ready for the machine god, don’t help them build it faster” is super simple?
I guess math very much isn’t like this. You can generally make at least some progress on a problem, but you can’t really make a problem worse except by giving bad advice to others working on it?
I think there’s something especially bad about AI, where people disagree a lot or are confused about what exactly the problems are, and have a lot of motivated cognition towards thinking “if they’re working on it, it’ll go better instead of worse.”
(i.e. if people were right that the main problem is misuse, going to work at a lab is more reasonable)
There’s something additionally hard about the existential stakes, where it’s hard to admit the problem is real if you don’t feel like you have some way of engaging with it.
E.g. Musk seems like mega-biased toward taking action, with bad results (responsible for like a quarter of frontier AI lab stuff, in some sense). The politician’s syllogism bites real hard when accelerating bad stuff is 1. convergent and 2. much easier than accelerating good stuff.
“Please read IABIED” is my first thought, but sounds like you’ve probably tried that? Maybe there should be a TLDR aimed at mathematicians? We could ask Yudkowsky to write one?
I think recommending a book just doesn’t work without a long preamble or some unusual social leverage. I’ve had better success recommending online articles, my current go-to is Scott’s book review of IABIED, which is still tl;dr and has a bunch of missing context.
The main thing a good first-intro has to do, honestly, is to be as short as possible while explaining instrumental convergence, orthogonality, and countering some representative examples of crazy moon logic that people come up with initially.
Have you tried AISafety.info’s intro sequence? There’s also a shorter stand-alone article, but that one is written for the average person.
Thanks for the tip! On first pass I think it is more measured and comprehensive but less fun-to-read than Scott’s review so I would likely not try that instead.
TBH, I thought mathematicians would prefer that.
I’d be curious if you (or anybody else) want to test to see if my intro is superior[1]. I weakly think it would be, but I’m pretty unconfident and there are good reasons for me to share my guide with people I know over other people’s guides regardless of objective merit.
I deliberately wrote the intro to avoid many mistakes I saw in other intros (eg too many extraneous details, appealing to overly complex analogies, shadowboxing insider objections, or “going meta” in an intro post).
eg by randomizing my intro vs Scott’s review to send to different people.