During the 2012-2014-ish era, I recall a fairly common statement from MIRI being “you might think AI alignment is a computer programming problem but it’s more of a math problem”, and there was some attempt to carve it into fun little chunk for Math people.
Then MIRI mostly gave up on solving it’s technical agenda in time. I’m not sure what the state of things are in terms of “what problems are still open and seem to matter?”.
Seems probably worth someone putting in legwork to organize the state of the problem into something comprehensible to the math world. I vaguely am thinking of @Max Harms maybe being a good fit for this?
(I’ve been “out of the game” for a couple years.) I would append to
you might think AI alignment is a computer programming problem but it’s more of a math problem
a follow-on:
you might think AI alignment is a math problem but it’s more of a philosophical / conceptual / mental-phenomenology problem
(I hesitate to use the word “phenomenology” because it will definitely be misunderstood by ~everyone, but there just isn’t a better word; an internal phrase that was floated was “core theory of mind”, and I mean to gesture at “mathematico-introspective investigation of core theory of mind”.) This is elaborated on (cryptically / elliptically, but you could read slowly & think) here: https://www.lesswrong.com/posts/TNQKFoWhAkLCB4Kt7/a-hermeneutic-net-for-agency
If you take all the alignment research, insofar as I’m aware of it, in the past 2 decades, and multiply that by 3, it still doesn’t even come close to solving alignment. This is very rough and grim, but it’s what I think.
I’m not sure I buy that most mathematicians will be out of a job. Maybe some of the number theorists, analysts, and combinatorialists? Or something, IDK, I’m making that up. (Based on what fields vaguely seem to have some niches for people doing mostly high-algebraicness work.) Or maybe they will be.
What should they do instead? IDK. Human intelligence amplification is still insanely underresearched, and can definitely use more brainpower. I’d be very happy to be connected with smart motivated scientists / mathematicians who may want to work on empowering humans!! My gmail for that would be: tsvibtcontact
To be clear, when writing this post, recruiting any mathematicians to alignment work was much more of a stretch goal.
The low-hanging fruit I’m trying to point to is that there are droves of academics currently considering going to work at frontier AI labs without having even heard the basic arguments for x-risk. I think it would be much much easier to convince a substantial fraction of them to do NOT THAT, than to direct them to do any active technical alignment work.
Having mulled this slightly more… I don’t actually have a very clear idea on what would work here (in the sense of “be memetically successful enough to reach a meaningful number of people.”)
I think there’s a moderately effortful job here of “figure out how to communicate about this in terms the math community will be motivated by and engage with” which can only really be done by someone in the math community.
I’m curious, when you imagine doing “the obvious thing”, what ends up being hard or not working? (presumably there’s a reason you wrote this post targeted at LW users than some other post targeted at math people?)
The “obvious thing” that I have done is state my views on doom to people when it comes up and point them to IABIED, and I think it has largely been successful in causing ~20 people I have personal relationships with to arrive at p(doom) > 5%.
The next “obvious thing” is to scale this up by a factor of 10 to people I’m only vaguely acquainted with, and to write more publicly in a mathematician-facing way as you suggested. These feel both costly, risky, and less effective in ways that the previous thing was not. I am still considering doing it, but need to sit down with my aversions for a while. Some examples: these conversations sometimes require me to be forceful/activist to keep people in my frame instead of sliding off into more comfortable topics, and they plausibly only work because I have significant personal rapport already with a colleague or student.
I would certainly prefer to live in a world where 1 in 20 mathematicians are already awake and I can just create common knowledge that the private conversation thing works, and then not do anything more effortful. As far as I know the only other academic mathematician who writes about Doom is Jacob Tsimerman, who just won the Fields Medal and left to do safety at OpenAI. He would certainly be a person who could do the broadcasting thing (and is doing it already, but in a less forceful way than I’m imagining).
(This is probably not helpful, but just want to note that I’m interested, in a “hobbyist” sense, in the problem of the general kind of confrontation you’re mentioning, under the moniker “confrontation-worthy empathy”; if you happen to want to chat about it I’m interested, e.g. to consider different ideas for making the process more wholesome and similar.)
I would be interested in having this conversation! It seems to me that the kind of confrontation I’m mentioning is a much easier version of the thing you’re pointing to? I tend to just point amateur levels of unconditional positive regard at people and under pre-conditions of existing trust conversations mostly go well. I imagine scaling up the operation requires more finesse, but even so it’s hard to imagine needing the amount of skill you describe there. I’ll message you.
Partly, I am kinda assuming people have a harder time wrapping their brain around “what not to do” vs “what to do”, and I think you’ll be more successful with that goal if you present a positive vision of how to orient to x-risk that speaks their language.
But, yeah that is a pretty different framing that would be suspicious if it didn’t look different at all from where I was pointing.
Partly, I am kinda assuming people have a harder time wrapping their brain around “what not to do” vs “what to do”
IDK, there being problems that are hard to fix but easy to make worse seems like a sort of common sense situation? Though I’m struggling to think of any good examples off the top of my head. Also just “we aren’t ready for the machine god, don’t help them build it faster” is super simple?
I guess math very much isn’t like this. You can generally make at least some progress on a problem, but you can’t really make a problem worse except by giving bad advice to others working on it?
I think there’s something especially bad about AI, where people disagree a lot or are confused about what exactly the problems are, and have a lot of motivated cognition towards thinking “if they’re working on it, it’ll go better instead of worse.”
(i.e. if people were right that the main problem is misuse, going to work at a lab is more reasonable)
There’s something additionally hard about the existential stakes, where it’s hard to admit the problem is real if you don’t feel like you have some way of engaging with it.
E.g. Musk seems like mega-biased toward taking action, with bad results (responsible for like a quarter of frontier AI lab stuff, in some sense). The politician’s syllogism bites real hard when accelerating bad stuff is 1. convergent and 2. much easier than accelerating good stuff.
“Please read IABIED” is my first thought, but sounds like you’ve probably tried that? Maybe there should be a TLDR aimed at mathematicians? We could ask Yudkowsky to write one?
I think recommending a book just doesn’t work without a long preamble or some unusual social leverage. I’ve had better success recommending online articles, my current go-to is Scott’s book review of IABIED, which is still tl;dr and has a bunch of missing context.
The main thing a good first-intro has to do, honestly, is to be as short as possible while explaining instrumental convergence, orthogonality, and countering some representative examples of crazy moon logic that people come up with initially.
Thanks for the tip! On first pass I think it is more measured and comprehensive but less fun-to-read than Scott’s review so I would likely not try that instead.
I’d be curious if you (or anybody else) want to test to seeif my intro is superior[1]. I weakly think it would be, but I’m pretty unconfident and there are good reasons for me to share my guide with people I know over other people’s guides regardless of objective merit.
I deliberately wrote the intro to avoid many mistakes I saw in other intros (eg too many extraneous details, appealing to overly complex analogies, shadowboxing insider objections, or “going meta” in an intro post).
My sense is many of these problems are either ill-defined or too hard to be tractable.
In certain fields like computability theory most problems are intractable, just because programs are very complicated and diverse objects that are hard to prove things about. Progress in such fields is made by working in the areas of the field where there is enough structure. Unfortunately proofs over programs have featured in agent foundations since the beginning (eg tiling agents).
As for ill-defined problems, ontology identification, embedded agency, and decision theory are full of them. Eg finding something that behaves like counterlogicals, which are nonexistent objects. Because they are nonexistent objects it requires philosophical progress to make a list of properties they need to satisfy. This doesn’t mean it’s impossible to make progress, but the problems need to be formalized in a way that are tractable and don’t lose all their relevance to AI.
A specific thing I’m curious about: there was a strategy of “turn some open alignment problems into nice little math problem to nerd-snipe people with.” Did that strategy accomplish anything? (i.e did it successfully get someone who would have bounced off the general conceptual arguments to engage enough to seem to eventually get the conceptual argument? And/or just produce obviously useful work)
During the 2012-2014-ish era, I recall a fairly common statement from MIRI being “you might think AI alignment is a computer programming problem but it’s more of a math problem”, and there was some attempt to carve it into fun little chunk for Math people.
Then MIRI mostly gave up on solving it’s technical agenda in time. I’m not sure what the state of things are in terms of “what problems are still open and seem to matter?”.
Seems probably worth someone putting in legwork to organize the state of the problem into something comprehensible to the math world. I vaguely am thinking of @Max Harms maybe being a good fit for this?
Curious for braindumps from @abramdemski @Scott Garrabrant, @TsviBT, @johnswentworth and @Alex_Altair on the state of thing heres.
(I’ve been “out of the game” for a couple years.) I would append to
a follow-on:
(I hesitate to use the word “phenomenology” because it will definitely be misunderstood by ~everyone, but there just isn’t a better word; an internal phrase that was floated was “core theory of mind”, and I mean to gesture at “mathematico-introspective investigation of core theory of mind”.) This is elaborated on (cryptically / elliptically, but you could read slowly & think) here: https://www.lesswrong.com/posts/TNQKFoWhAkLCB4Kt7/a-hermeneutic-net-for-agency
In this context, the long and short of it is, I don’t think you can feasibly point people at the parts of the problem that matter.
I’m just such a special genius that only I could possibly understand the real alignment problemFor some reason people don’t seem to be inclined to plant difficult questions and let them grow over time, sit with unresolved questions, look at big things and keep their bigness firmly in mind while also somehow making cumulative progress, overhaul key elements or at least keep them firmly provisional, keep staring outside of the streelight’s glow for years, and so forth. (Something something Grothendieck something something Peter Scholze something something staring at conceptual foundations.) Or maybe it’s the thing about security mindset (Yudkowsky), or overconfidence (Dai), etc. Or maybe we’re just not high g enough. Or all of the above or something else, IDK. But again, it just doesn’t seem to work to point people at the problem. Or maybe I’m deluded, who could know.If you take all the alignment research, insofar as I’m aware of it, in the past 2 decades, and multiply that by 3, it still doesn’t even come close to solving alignment. This is very rough and grim, but it’s what I think.
I’m not sure I buy that most mathematicians will be out of a job. Maybe some of the number theorists, analysts, and combinatorialists? Or something, IDK, I’m making that up. (Based on what fields vaguely seem to have some niches for people doing mostly high-algebraicness work.) Or maybe they will be.
What should they do instead? IDK. Human intelligence amplification is still insanely underresearched, and can definitely use more brainpower. I’d be very happy to be connected with smart motivated scientists / mathematicians who may want to work on empowering humans!! My gmail for that would be: tsvibtcontact
To be clear, when writing this post, recruiting any mathematicians to alignment work was much more of a stretch goal.
The low-hanging fruit I’m trying to point to is that there are droves of academics currently considering going to work at frontier AI labs without having even heard the basic arguments for x-risk. I think it would be much much easier to convince a substantial fraction of them to do NOT THAT, than to direct them to do any active technical alignment work.
Having mulled this slightly more… I don’t actually have a very clear idea on what would work here (in the sense of “be memetically successful enough to reach a meaningful number of people.”)
I think there’s a moderately effortful job here of “figure out how to communicate about this in terms the math community will be motivated by and engage with” which can only really be done by someone in the math community.
I’m curious, when you imagine doing “the obvious thing”, what ends up being hard or not working? (presumably there’s a reason you wrote this post targeted at LW users than some other post targeted at math people?)
The “obvious thing” that I have done is state my views on doom to people when it comes up and point them to IABIED, and I think it has largely been successful in causing ~20 people I have personal relationships with to arrive at p(doom) > 5%.
The next “obvious thing” is to scale this up by a factor of 10 to people I’m only vaguely acquainted with, and to write more publicly in a mathematician-facing way as you suggested. These feel both costly, risky, and less effective in ways that the previous thing was not. I am still considering doing it, but need to sit down with my aversions for a while. Some examples: these conversations sometimes require me to be forceful/activist to keep people in my frame instead of sliding off into more comfortable topics, and they plausibly only work because I have significant personal rapport already with a colleague or student.
I would certainly prefer to live in a world where 1 in 20 mathematicians are already awake and I can just create common knowledge that the private conversation thing works, and then not do anything more effortful. As far as I know the only other academic mathematician who writes about Doom is Jacob Tsimerman, who just won the Fields Medal and left to do safety at OpenAI. He would certainly be a person who could do the broadcasting thing (and is doing it already, but in a less forceful way than I’m imagining).
(This is probably not helpful, but just want to note that I’m interested, in a “hobbyist” sense, in the problem of the general kind of confrontation you’re mentioning, under the moniker “confrontation-worthy empathy”; if you happen to want to chat about it I’m interested, e.g. to consider different ideas for making the process more wholesome and similar.)
I would be interested in having this conversation! It seems to me that the kind of confrontation I’m mentioning is a much easier version of the thing you’re pointing to? I tend to just point amateur levels of unconditional positive regard at people and under pre-conditions of existing trust conversations mostly go well. I imagine scaling up the operation requires more finesse, but even so it’s hard to imagine needing the amount of skill you describe there. I’ll message you.
Yeah, I think it should be much less hard for people not already embedded in working on AI; just seems kinda related.
Yeah fair.
Partly, I am kinda assuming people have a harder time wrapping their brain around “what not to do” vs “what to do”, and I think you’ll be more successful with that goal if you present a positive vision of how to orient to x-risk that speaks their language.
But, yeah that is a pretty different framing that would be suspicious if it didn’t look different at all from where I was pointing.
IDK, there being problems that are hard to fix but easy to make worse seems like a sort of common sense situation? Though I’m struggling to think of any good examples off the top of my head. Also just “we aren’t ready for the machine god, don’t help them build it faster” is super simple?
I guess math very much isn’t like this. You can generally make at least some progress on a problem, but you can’t really make a problem worse except by giving bad advice to others working on it?
I think there’s something especially bad about AI, where people disagree a lot or are confused about what exactly the problems are, and have a lot of motivated cognition towards thinking “if they’re working on it, it’ll go better instead of worse.”
(i.e. if people were right that the main problem is misuse, going to work at a lab is more reasonable)
There’s something additionally hard about the existential stakes, where it’s hard to admit the problem is real if you don’t feel like you have some way of engaging with it.
E.g. Musk seems like mega-biased toward taking action, with bad results (responsible for like a quarter of frontier AI lab stuff, in some sense). The politician’s syllogism bites real hard when accelerating bad stuff is 1. convergent and 2. much easier than accelerating good stuff.
“Please read IABIED” is my first thought, but sounds like you’ve probably tried that? Maybe there should be a TLDR aimed at mathematicians? We could ask Yudkowsky to write one?
I think recommending a book just doesn’t work without a long preamble or some unusual social leverage. I’ve had better success recommending online articles, my current go-to is Scott’s book review of IABIED, which is still tl;dr and has a bunch of missing context.
The main thing a good first-intro has to do, honestly, is to be as short as possible while explaining instrumental convergence, orthogonality, and countering some representative examples of crazy moon logic that people come up with initially.
Have you tried AISafety.info’s intro sequence? There’s also a shorter stand-alone article, but that one is written for the average person.
Thanks for the tip! On first pass I think it is more measured and comprehensive but less fun-to-read than Scott’s review so I would likely not try that instead.
TBH, I thought mathematicians would prefer that.
I’d be curious if you (or anybody else) want to test to see if my intro is superior[1]. I weakly think it would be, but I’m pretty unconfident and there are good reasons for me to share my guide with people I know over other people’s guides regardless of objective merit.
I deliberately wrote the intro to avoid many mistakes I saw in other intros (eg too many extraneous details, appealing to overly complex analogies, shadowboxing insider objections, or “going meta” in an intro post).
eg by randomizing my intro vs Scott’s review to send to different people.
My sense is many of these problems are either ill-defined or too hard to be tractable.
In certain fields like computability theory most problems are intractable, just because programs are very complicated and diverse objects that are hard to prove things about. Progress in such fields is made by working in the areas of the field where there is enough structure. Unfortunately proofs over programs have featured in agent foundations since the beginning (eg tiling agents).
As for ill-defined problems, ontology identification, embedded agency, and decision theory are full of them. Eg finding something that behaves like counterlogicals, which are nonexistent objects. Because they are nonexistent objects it requires philosophical progress to make a list of properties they need to satisfy. This doesn’t mean it’s impossible to make progress, but the problems need to be formalized in a way that are tractable and don’t lose all their relevance to AI.
I’ll also point to the people and agenda at Resolution: https://resolution.org/launch
A specific thing I’m curious about: there was a strategy of “turn some open alignment problems into nice little math problem to nerd-snipe people with.” Did that strategy accomplish anything? (i.e did it successfully get someone who would have bounced off the general conceptual arguments to engage enough to seem to eventually get the conceptual argument? And/or just produce obviously useful work)