I think this imagines a world where we continue current trend lines for a few more years, but don’t get AI that navigates the world and executes clever general plans like humans do. LLMs just get a little bit smarter, and a lot more work is put into making them good at specific domains (similar to what’s been done for coding).
(But AI never actually acts autonomously to compete with humanity in this imagined future. Either it never gets easy to get an AI to act like a clever goal-directed agent that navigates the real world, or it does become easy but somehow it never happens by accident and also nobody does it on purpose.)
So we can have “even ‘ASI’ may be bad at [...] moral philosophy and social planning.”, and in this imagined scenario it doesn’t mean “the AI is amoral and we’re fucked,” but “the AI is amoral and this is a moderate inconvenience.” This thing being referred to as “ASI” is just a few years linear extrapolation of current LLMs—it’s superhuman at, say, managing factories only because humans have poured millions of dollars into building factory-management RL environments, and millions more dollars into training a big base model on them. It doesn’t learn like humans do, and very especially it doesn’t navigate the world in general like humans do—it’s missing some of the key skills to do so, even though it has lots of skills for factory management.
Amoral ‘ASI’ of this sort can’t do anything complicated in the world without human help at many key steps. Not just menial help, actual help in the decision-making process. So when bad stuff happens, there’s a bunch of convenient humans involved to blame.
Anyhow, yeah, I agree, building AI “bad at moral philosophy” is bad even in near-term extrapolations where it’s just running factories amorally and causing societal upheaval and making people sad. It’s just that also, there’s this thing where once the AI can do complicated stuff in the real world by itself, it being amoral is extra bad.
People who have never read Kuhn or Lakatos or Feyerabend
I’ve read some good excerpts from Against Method, but on the whole I think mentioning Feyrabend is a red flag. If this person were serious, they might have brought up Li and Vitanyi instead.
“even ‘ASI’ may be bad at [...] moral philosophy and social planning.”
To be clear, this thought was mine, I don’t think it appeared in the linked post. There are people thinking about this sort of thing at Redwood (or rather, conceptual reasoning, which I claim includes moral philosophy and social planning among other things). The most up-to-date link I can point to is Current AIs seem pretty misaligned to me; also, this post gives some idea of a mitigation.
I could imagine a world where AI “truly generalizes” enough to run a company, but its attempts at moral philosophy and social planning are still mostly slop. Successfully running a company is in fact measurable, even if it involves pretty long timescales (just maximize revenue and/or the stock price). But for other domains, we have no good way to measure success. I feel very uncertain about the relationship between these two types of capabilities.
Social planning is very measurable. Could you achieve goals in a complicated social situation?
I agree moral philosophy is inherently different—when we say we want an AI to be “good” at moral reasoning, we’re self-consciously referencing our own vague human standards for good moral reasoning. The problem an AI has to solve to get good at moral reasoning is not just about induction, but about communication, and we might build AI that does the former but not the latter.
Thank you for the summary.
I think this imagines a world where we continue current trend lines for a few more years, but don’t get AI that navigates the world and executes clever general plans like humans do. LLMs just get a little bit smarter, and a lot more work is put into making them good at specific domains (similar to what’s been done for coding).
(But AI never actually acts autonomously to compete with humanity in this imagined future. Either it never gets easy to get an AI to act like a clever goal-directed agent that navigates the real world, or it does become easy but somehow it never happens by accident and also nobody does it on purpose.)
So we can have “even ‘ASI’ may be bad at [...] moral philosophy and social planning.”, and in this imagined scenario it doesn’t mean “the AI is amoral and we’re fucked,” but “the AI is amoral and this is a moderate inconvenience.” This thing being referred to as “ASI” is just a few years linear extrapolation of current LLMs—it’s superhuman at, say, managing factories only because humans have poured millions of dollars into building factory-management RL environments, and millions more dollars into training a big base model on them. It doesn’t learn like humans do, and very especially it doesn’t navigate the world in general like humans do—it’s missing some of the key skills to do so, even though it has lots of skills for factory management.
Amoral ‘ASI’ of this sort can’t do anything complicated in the world without human help at many key steps. Not just menial help, actual help in the decision-making process. So when bad stuff happens, there’s a bunch of convenient humans involved to blame.
Anyhow, yeah, I agree, building AI “bad at moral philosophy” is bad even in near-term extrapolations where it’s just running factories amorally and causing societal upheaval and making people sad. It’s just that also, there’s this thing where once the AI can do complicated stuff in the real world by itself, it being amoral is extra bad.
I’ve read some good excerpts from Against Method, but on the whole I think mentioning Feyrabend is a red flag. If this person were serious, they might have brought up Li and Vitanyi instead.
To be clear, this thought was mine, I don’t think it appeared in the linked post. There are people thinking about this sort of thing at Redwood (or rather, conceptual reasoning, which I claim includes moral philosophy and social planning among other things). The most up-to-date link I can point to is Current AIs seem pretty misaligned to me; also, this post gives some idea of a mitigation.
I could imagine a world where AI “truly generalizes” enough to run a company, but its attempts at moral philosophy and social planning are still mostly slop. Successfully running a company is in fact measurable, even if it involves pretty long timescales (just maximize revenue and/or the stock price). But for other domains, we have no good way to measure success. I feel very uncertain about the relationship between these two types of capabilities.
Social planning is very measurable. Could you achieve goals in a complicated social situation?
I agree moral philosophy is inherently different—when we say we want an AI to be “good” at moral reasoning, we’re self-consciously referencing our own vague human standards for good moral reasoning. The problem an AI has to solve to get good at moral reasoning is not just about induction, but about communication, and we might build AI that does the former but not the latter.
I meant the kind of social planning described on this (SEO slop) webpage. Basically “how do we develop prosocial government policies?”