Context: I’m a mathematics professor in the Netherlands, and one of the coauthors of the Leiden Declaration on Artificial Intelligence and Mathematics. I think this article is was a nice summary of the current situation. My impression from talking to my colleagues is not that there is a lack of concern; people have not necessarily thought through the details, but the idea of existential risk from AI is taken seriously. What is missing is any clear idea of what to do about it; people feel powerless. Near the end of your essay you wrote:
“Many people say that alignment is primarily a mathematics problem, so I think we have a lot to offer.”
I’m curious if you can make that more concrete? I expect that if someone stated a reasonably concrete mathematical conjecture and said “solving this would really help with existential risk” there would be plenty of people happy to devote effort to it. But my own (perhaps very wrong) impression of the field it that the questions are more like “come up with a solution to inner alignment” which is too vague for most mathematicians to make a start on. There are some very mathematical alignment programmes out there (e.g. that of Vanessa Kosoy); do you think that having lots of people pile on a programme like that would help?
Somehow it would be great to have something like the Langland’s Programme for AI alignment; it can be hard and incompletely specified, but it gives a convincing direction, is concrete enough to get started on, and it’s clear that if you succeed you have really made progress on the problem. At the moment that seems to me to be missing, but I’d love to be corrected.
Hi David! Thanks for your input and for the work on the Leiden Declaration. My essay was intentionally vague on what to do, as I don’t strongly believe in a particular approach that is likely to help, and also have not worked deeply enough in technical alignment to have strong opinions.
To be honest, one of the main goals of this essay was an implicit call for inaction: “Don’t go work for an AI company without sitting down and thinking about x-risk (but maybe do if you have a concrete theory of change after thinking about it, like Tsimerman seems to).”
Some obvious places for mathematicians to start looking for concrete directions are ARC and Lionel Levine’s work. Many major universities already have extant AI Safety groups, so senior mathematicians can probably pick up some low-hanging fruit just by getting affiliated to lend credibility, and asking to be looped in on the mathematical parts of their research.
My current conjecture (lightly held) is that senior mathematicians concerned about x-risk can have the most immediate impact not by trying to solve technical problems, as we are conditioned to do, but by throwing our weight around in universities and public discourse until we reach a critical mass. I understand that many individual mathematicians are becoming aware of the case for x-risk, but there is almost 0 public discourse or common knowledge about it. I think we could very easily completely shift the discourse if ~20 mathematicians wrote their own version of my essay, or there was a version of the Leiden Declaration centered around alignment concerns.
Thinking about this a bit more, I suspect you’re experiencing some kind of Different Worlds effect, several other people I’ve talked to report that every mathematician around them is completely dismissive of x-risk...
I was around at the workshop leading to the Leiden Declaration and was actively trying to talk to people about x-risk. I can confirm that 9⁄10 people I talked to seemed either dismissive or incredulous, with a few people being quite open to the concerns. There was, however, far from enough critical mass for this to gain any traction in the broader discussion and instead a lot of space was taken up by more proximate questions that somehow seemed to assume AI progress would stall at current levels. I can imagine this changed more recently, and separately that people the Netherlands tend to be a bit more forward thinking on these matters anyway (just based on my general impression).
Context: I’m a mathematics professor in the Netherlands, and one of the coauthors of the Leiden Declaration on Artificial Intelligence and Mathematics. I think this article is was a nice summary of the current situation. My impression from talking to my colleagues is not that there is a lack of concern; people have not necessarily thought through the details, but the idea of existential risk from AI is taken seriously. What is missing is any clear idea of what to do about it; people feel powerless. Near the end of your essay you wrote:
“Many people say that alignment is primarily a mathematics problem, so I think we have a lot to offer.”
I’m curious if you can make that more concrete? I expect that if someone stated a reasonably concrete mathematical conjecture and said “solving this would really help with existential risk” there would be plenty of people happy to devote effort to it. But my own (perhaps very wrong) impression of the field it that the questions are more like “come up with a solution to inner alignment” which is too vague for most mathematicians to make a start on. There are some very mathematical alignment programmes out there (e.g. that of Vanessa Kosoy); do you think that having lots of people pile on a programme like that would help?
Somehow it would be great to have something like the Langland’s Programme for AI alignment; it can be hard and incompletely specified, but it gives a convincing direction, is concrete enough to get started on, and it’s clear that if you succeed you have really made progress on the problem. At the moment that seems to me to be missing, but I’d love to be corrected.
Hi David! Thanks for your input and for the work on the Leiden Declaration. My essay was intentionally vague on what to do, as I don’t strongly believe in a particular approach that is likely to help, and also have not worked deeply enough in technical alignment to have strong opinions.
To be honest, one of the main goals of this essay was an implicit call for inaction: “Don’t go work for an AI company without sitting down and thinking about x-risk (but maybe do if you have a concrete theory of change after thinking about it, like Tsimerman seems to).”
Some obvious places for mathematicians to start looking for concrete directions are ARC and Lionel Levine’s work. Many major universities already have extant AI Safety groups, so senior mathematicians can probably pick up some low-hanging fruit just by getting affiliated to lend credibility, and asking to be looped in on the mathematical parts of their research.
My current conjecture (lightly held) is that senior mathematicians concerned about x-risk can have the most immediate impact not by trying to solve technical problems, as we are conditioned to do, but by throwing our weight around in universities and public discourse until we reach a critical mass. I understand that many individual mathematicians are becoming aware of the case for x-risk, but there is almost 0 public discourse or common knowledge about it. I think we could very easily completely shift the discourse if ~20 mathematicians wrote their own version of my essay, or there was a version of the Leiden Declaration centered around alignment concerns.
I am personally a fan of Resolution’s program …
https://www.lesswrong.com/posts/AP7YDke5jjY4v3X9Z/resolution-fka-sequent-scale-and-automation-for-higher
much of which is still listed here …
https://timaeus.co/projects
there is also …
https://www.lesswrong.com/posts/KfkpgXdgRheSRWDy8/a-list-of-45-mech-interp-project-ideas-from-apollo-research
(a little over 2 years old now) and …
https://www.lesswrong.com/posts/mtGpdtDdmkRC3ZBuz/list-of-lists-of-project-ideas-in-ai-safety
which is not quite a year old and a little more diverse.
oh my gosh yes. The woods are lovely, dark and deep, But I have promises to keep, And miles to go before I sleep, etc.
Thanks, this motivated me to search around a bit more, the situation is better than last time I looked; I’ve tarted reading Lionel Levine’s recent paper [https://github.com/lionellevine/MAIS/blob/main/papers/P1/MAIS-P1.pdf].
Thinking about this a bit more, I suspect you’re experiencing some kind of Different Worlds effect, several other people I’ve talked to report that every mathematician around them is completely dismissive of x-risk...
I was around at the workshop leading to the Leiden Declaration and was actively trying to talk to people about x-risk. I can confirm that 9⁄10 people I talked to seemed either dismissive or incredulous, with a few people being quite open to the concerns.
There was, however, far from enough critical mass for this to gain any traction in the broader discussion and instead a lot of space was taken up by more proximate questions that somehow seemed to assume AI progress would stall at current levels.
I can imagine this changed more recently, and separately that people the Netherlands tend to be a bit more forward thinking on these matters anyway (just based on my general impression).