These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”: viewing an outcome as so inevitable that it doesn’t matter much if you contribute to it, in a way which leads you to become a significant force pushing the world further and faster towards the “inevitable” outcome.
I agree that this pattern is real and that EAs and AI safety people such as myself have engaged in it too much historically. But how far do you think we should go in avoiding this failure mode? Should we refuse to use ChatGPT? Should we refuse to talk publicly about superintelligence, and instead dismiss it all as hype, in the hope of bursting the AI bubble? (There are probably many thousands of people currently taking this strategy so it’s not just a hypothetical!) My current answer to both questions is no and I expect you’ll agree, so there’s a question of where to draw the line. I’m interested in ideas for principled and/or historically validated answers.
My knee-jerk reaction here is “have you tried asking a five year old?”. Spelling it out in more words: for most of these questions there is an obvious Right Thing To Do. Say true things, spread true important things, don’t lie or strategically hide your views, don’t accelerate capabilities, don’t found capabilities companies, don’t bullshit. Just normal five-year-old level ethics.
And look, I am explicitly on the record saying Value Judgements Are Usually Bullshit and Human Values ≠ Goodness and just generally opposing mainstream morality. But even I can look at the track record of EA and AI Safety and say “look guys, you have not been outperforming the Obvious Right Thing To Do with respect to AI, you should stop with the gigabrain takes and just do the Obvious Right Things”.
(Also, my opposition to mainstream morality diverges from normal five-year-old level ethics in basically the opposite direction from gigabrain EA/AI safety takes. Compared to the five-year-old, I lean even more heavily toward saying true important things, and even less toward leading Molochian parades in exchange for nominal authority.)
The two paragraphs after the one you quoted are intended to give some guidance on this. Some elaboration on key points, moving from most to least actionable (many of which I expect you’re aware of, but this is just what I have off the top of my head):
Safety people keep dramatically underrating how much more seriously they take superintelligence than everyone else. Of course they have a bunch of evidence that other people don’t take superintelligence seriously (e.g. from constantly being told that caring about alignment is silly)—they just don’t cross-apply it to questions like “how much of the work at OpenAI or DeepMind is actually oriented towards building superintelligence”. A reasonable approximation to work from is that no capabilities researcher takes superintelligence seriously by rationalist standards, otherwise they wouldn’t be a capabilities researcher (unless they’re exceptionally power-seeking). Even the people who invented reasoning models don’t take AGI seriously (e.g. when I asked the team lead about what he thought AI capabilities would be like in two years, he sort of shrugged and told me he wasn’t really thinking about that). To some extent, this is evidence that you can make a bunch of progress by not taking superintelligence seriously; but it’s also evidence that they’ll do less out-of-distribution reasoning about how to build it than you might think.
A more general reasonable heuristic is that you should think of these things as varying by many orders of magnitude. For example, how seriously you take an idea could be measured by how much social pressure it would take to make you stop considering it. Some people shape their beliefs around even very small negative feedback (like someone giving them the side-eye); others are willing to stand up to billions calling them evil. Similarly, “amount of risk you’re personally willing to take” varies by orders of magnitude; as does “how adversarially are the US and China treating each other vis a vis AI”.
You should think about your actions in part as a process of upholding and maintaining norms. The ways you should choose to act depend in significant part on which coalitions you want to throw your weight behind. (Scott Alexander discusses some similar dynamics here—I don’t really endorse the post, but it’s at least trying to reason through such questions.) In general, you should err significantly on the side of not ruling out norms you think would be good, even if they seem overly ambitious. For example, it doesn’t seem crazy to me that a norm of “no rationalist should work at OpenAI” could have been enforced when it was first founded.
In principle this is related to FDT but doing any actual FDT reasoning is so complicated that you should in practice just think more about sociology instead.
Kantianism is an extreme version of this where you in some sense assume that you’re in a coalition with everyone. This is going too far.
As I say in a later paragraph: “even given these conceptual errors, people wouldn’t jump down nearly as many slippery slopes if they weren’t driven by strong emotional instincts”. My sequence on Replacing Fear is my best attempt to articulate how to do better on this axis. But it seems like various therapeutic modalities (especially the more somatic/embodied ones) work better than any cerebral intervention. The interventions that have been most helpful for me were circling, shrooms, bodywork, and Sleepawake. (To be clear, I think that hippies are missing something important about virtue/integrity, so I don’t recommend adopting their overall worldview, but their tools are very useful regardless.)
In principle this is related to FDT but doing any actual FDT reasoning is so complicated
I’ve seen you say something like this before, FWIW I don’t really agree that normal social hyperstition-style reasoning is FDT reasoning (EDIT: in the way that you mean it here) (e.g. this post of mine). (it can be used to give it additional justification, but it gets extremely confusing and metaphysical and it’s unclear to me what the direction of the update even is (cf my other posts in that sequence)).
There is a high-level philosophical consonance which I agree is important, but my guess is that’s not what you mean there.
EDIT: On reflection, the secondary claims here were not quite right, but I stand by the primary claim (although to be clear, the linked post doesn’t justify it specifically). The more precise thing to say is that if you think of FDT reasoning as “complicated”, my guess is you have a wrong conception of it (a naive TDT-style one) that doesn’t fully understand the reason why hyperstition-style reasoning works (which is updatelessness, not logical counterfactuals). But this isn’t a proper justification of that, I just wanted to be clear about what I’m saying.
A reasonable approximation to work from is that no capabilities researcher takes superintelligence seriously by rationalist standards, otherwise they wouldn’t be a capabilities researcher (unless they’re exceptionally power-seeking).
Is that even a good power-seeking strategy? It seems like maybe if you can produce technical contributions to or play politics well enough to gain substantial leverage at an AGI lab, this could be a good play? But I’m kind of surprised if it’s the best strategy for most AGI-pilled power-seekers.
Maybe you mean it in a more mundane way—that they get paid a lot of money.
This is the best initial attempt I’ve seen at answering this question (or something like it, anyway). I don’t think it has a clear answer to many of the thorny cases though.
I agree that this pattern is real and that EAs and AI safety people such as myself have engaged in it too much historically. But how far do you think we should go in avoiding this failure mode? Should we refuse to use ChatGPT? Should we refuse to talk publicly about superintelligence, and instead dismiss it all as hype, in the hope of bursting the AI bubble? (There are probably many thousands of people currently taking this strategy so it’s not just a hypothetical!) My current answer to both questions is no and I expect you’ll agree, so there’s a question of where to draw the line. I’m interested in ideas for principled and/or historically validated answers.
My knee-jerk reaction here is “have you tried asking a five year old?”. Spelling it out in more words: for most of these questions there is an obvious Right Thing To Do. Say true things, spread true important things, don’t lie or strategically hide your views, don’t accelerate capabilities, don’t found capabilities companies, don’t bullshit. Just normal five-year-old level ethics.
And look, I am explicitly on the record saying Value Judgements Are Usually Bullshit and Human Values ≠ Goodness and just generally opposing mainstream morality. But even I can look at the track record of EA and AI Safety and say “look guys, you have not been outperforming the Obvious Right Thing To Do with respect to AI, you should stop with the gigabrain takes and just do the Obvious Right Things”.
(Also, my opposition to mainstream morality diverges from normal five-year-old level ethics in basically the opposite direction from gigabrain EA/AI safety takes. Compared to the five-year-old, I lean even more heavily toward saying true important things, and even less toward leading Molochian parades in exchange for nominal authority.)
The two paragraphs after the one you quoted are intended to give some guidance on this. Some elaboration on key points, moving from most to least actionable (many of which I expect you’re aware of, but this is just what I have off the top of my head):
Safety people keep dramatically underrating how much more seriously they take superintelligence than everyone else. Of course they have a bunch of evidence that other people don’t take superintelligence seriously (e.g. from constantly being told that caring about alignment is silly)—they just don’t cross-apply it to questions like “how much of the work at OpenAI or DeepMind is actually oriented towards building superintelligence”. A reasonable approximation to work from is that no capabilities researcher takes superintelligence seriously by rationalist standards, otherwise they wouldn’t be a capabilities researcher (unless they’re exceptionally power-seeking). Even the people who invented reasoning models don’t take AGI seriously (e.g. when I asked the team lead about what he thought AI capabilities would be like in two years, he sort of shrugged and told me he wasn’t really thinking about that). To some extent, this is evidence that you can make a bunch of progress by not taking superintelligence seriously; but it’s also evidence that they’ll do less out-of-distribution reasoning about how to build it than you might think.
A more general reasonable heuristic is that you should think of these things as varying by many orders of magnitude. For example, how seriously you take an idea could be measured by how much social pressure it would take to make you stop considering it. Some people shape their beliefs around even very small negative feedback (like someone giving them the side-eye); others are willing to stand up to billions calling them evil. Similarly, “amount of risk you’re personally willing to take” varies by orders of magnitude; as does “how adversarially are the US and China treating each other vis a vis AI”.
You should think about your actions in part as a process of upholding and maintaining norms. The ways you should choose to act depend in significant part on which coalitions you want to throw your weight behind. (Scott Alexander discusses some similar dynamics here—I don’t really endorse the post, but it’s at least trying to reason through such questions.) In general, you should err significantly on the side of not ruling out norms you think would be good, even if they seem overly ambitious. For example, it doesn’t seem crazy to me that a norm of “no rationalist should work at OpenAI” could have been enforced when it was first founded.
In principle this is related to FDT but doing any actual FDT reasoning is so complicated that you should in practice just think more about sociology instead.
Kantianism is an extreme version of this where you in some sense assume that you’re in a coalition with everyone. This is going too far.
As I say in a later paragraph: “even given these conceptual errors, people wouldn’t jump down nearly as many slippery slopes if they weren’t driven by strong emotional instincts”. My sequence on Replacing Fear is my best attempt to articulate how to do better on this axis. But it seems like various therapeutic modalities (especially the more somatic/embodied ones) work better than any cerebral intervention. The interventions that have been most helpful for me were circling, shrooms, bodywork, and Sleepawake. (To be clear, I think that hippies are missing something important about virtue/integrity, so I don’t recommend adopting their overall worldview, but their tools are very useful regardless.)
I’ve seen you say something like this before, FWIW I don’t really agree that normal social hyperstition-style reasoning is FDT reasoning (EDIT: in the way that you mean it here) (e.g. this post of mine). (
it can be used to give it additional justification, but it gets extremely confusing and metaphysical and it’s unclear to me what the direction of the update even is (cf my other posts in that sequence)).There is a high-level philosophical consonance which I agree is important, but my guess is that’s not what you mean there.EDIT: On reflection, the secondary claims here were not quite right, but I stand by the primary claim (although to be clear, the linked post doesn’t justify it specifically). The more precise thing to say is that if you think of FDT reasoning as “complicated”, my guess is you have a wrong conception of it (a naive TDT-style one) that doesn’t fully understand the reason why hyperstition-style reasoning works (which is updatelessness, not logical counterfactuals). But this isn’t a proper justification of that, I just wanted to be clear about what I’m saying.
Is that even a good power-seeking strategy? It seems like maybe if you can produce technical contributions to or play politics well enough to gain substantial leverage at an AGI lab, this could be a good play? But I’m kind of surprised if it’s the best strategy for most AGI-pilled power-seekers.
Maybe you mean it in a more mundane way—that they get paid a lot of money.
This is the best initial attempt I’ve seen at answering this question (or something like it, anyway). I don’t think it has a clear answer to many of the thorny cases though.