Some other reasons why a more engineering mindset was adopted in AI alignment as opposed to an idealized insight focused path that is more positive than the reasons you brought up is:
A lot of alignment/control agendas aren’t about being robust to arbitrarily capable models, but rather models of more bounded capabilities. The AI control agenda is easily the best example of this, where it happily admits that most of the methods here wouldn’t work if AIs got good enough at steganographic communication/neuralese to bypass control measures.
This is because we don’t actually need to solve the problem of AI alignment ourselves, and if we can automate AI alignment research, then in the absence of blockers, we wouldn’t need to do any work upfront (Of course, there’s a large and continuing debate on what blockers exist, and how hard they are to solve.)
The ‘fast takeoffs’ described in AI 2027/AI 2040 are months long takeoffs, which is near the upper end of the range of expectations of early-2010s LessWrong (I’d say 60th-70th percentile at least).
Under slower takeoffs, engineering mindset is fine, because you have more hopes on fixing the problem iteratively, and this is due to the fact that you can assume that AI capabilities are more bounded than thought.
This doesn’t mean engineering mindset is better than scientific mindset, but it does mean that engineering mindset isn’t as bad as you say.
And we should also admit that the research program you propose (agent foundations) has been underwhelming at best, and even though problems like the genie knows, but does not care has unfortunately been more accurate than we thought (mostly because we scaled up RL again, and are likely to keep doing this because of incentives), solutions from agent foundations have been lacking.
So in conclusion, while I could see engineering mindset being worse than scientific mindset, I also think that this is less obvious of a conclusion than you do (assuming the claim of engineering mindset being worse than scientific mindset is correct), and I wanted to present a partial positive case of engineering mindset to offset potential biases.
Science is partly a coupled process between engineering and theory, it is also hard to do causal inference from what counts as theory and what counts as practice.
Would you say that the discovery of the higgs boson was something that happened as a consequence through theoretical or practical physics?
What about deception, inner misalignment, general interpretability methods, if you trace their intelluectual lineage where do they come from? They’re not fully agent foundations but if you compare if they’re more ML based or coming from the larger space of theoretical alignment research I would attribute more causal influence to theoretical AI Safety. This is what agent foundations was up to like 3 years ago!
Or what do we mean by agent foundations here? What would you draw the boundaries around? Is it maybe better to use the word Theoretical AI Safety research?
Under slower takeoffs, engineering mindset is fine, because you have more hopes on fixing the problem iteratively, and this is due to the fact that you can assume that AI capabilities are more bounded than thought
Given specific assumptions about scientific progress where we can iteratively improve it and it is clear how we would even aim for it in the first place.
How much is this field biology and how much is it computer science? If it is computer science then it is bloody complicated combinatorial optimisation and dynamic programming. This field is in it’s philosophical nature closer to studying growing systems not engineered systems.
I would have much less beef with the engineering mindset if it hadn’t accelerated capabilities so much, as I’ll explain in the next post (and also if it weren’t so related to conceptually confused ML research, as I’ll explain in the post after that).
My main issue is The Counterfactual Quiet AGI Timeline. Scaling laws of neural nets had capabilities become more like treasures waiting for the right amount of compute. Once anyone invested the compute, the treasure would be his and everyone would rush for similar treasures, until one of them summons demons...
I don’t yet have a situation-global opinion here, but “once anyone invested the compute, the treasure would be theirs” doesn’t seem like a valid transition. The primary constraint in connecting abstract research to economic applications is generalist-capable domain-experts being motivated to bridge the gap from theory to something a VC or customer can understand.
While you can’t stay still forever relying on the absence of that force, a small non-prestigious community uniquely positioned to work in a field deciding not to seek investments and capability-research would certainly have further delayed the viability of the treasure, in turn supporting non-capabilities-dependent alignment-contribution tactics like intelligence augmentation, foundational research, philosophical research, or prestige-building from non-X-risk charitable efforts to acquire credibility for X-risk-mitigation.
There are valid replacement-reasons that could slot in the same place in your presented reasoning, including-but-not-limited-to longer timeline-estimates making personal-judgement-of-influencers more important than timeline-reductions, reason to believe some specific other party was already working on making transformer-based autonomously-operating AI economically viable at the time, expectation that capabilities could meaningfully help with other X-risks before they became a concern themselves, etc. I’m not sure which specifically of those you would agree with, though.
Some other reasons why a more engineering mindset was adopted in AI alignment as opposed to an idealized insight focused path that is more positive than the reasons you brought up is:
A lot of alignment/control agendas aren’t about being robust to arbitrarily capable models, but rather models of more bounded capabilities. The AI control agenda is easily the best example of this, where it happily admits that most of the methods here wouldn’t work if AIs got good enough at steganographic communication/neuralese to bypass control measures.
This is because we don’t actually need to solve the problem of AI alignment ourselves, and if we can automate AI alignment research, then in the absence of blockers, we wouldn’t need to do any work upfront (Of course, there’s a large and continuing debate on what blockers exist, and how hard they are to solve.)
We have gotten evidence for slower takeoffs than some early 2010s writings placed serious probability on. You’ve actually criticized AI 2040 in the past for focusing too much on fast takeoffs, and that is admittedly a fair enough criticism, but it is worth noting 2 things:
The ‘fast takeoffs’ described in AI 2027/AI 2040 are months long takeoffs, which is near the upper end of the range of expectations of early-2010s LessWrong (I’d say 60th-70th percentile at least).
Under slower takeoffs, engineering mindset is fine, because you have more hopes on fixing the problem iteratively, and this is due to the fact that you can assume that AI capabilities are more bounded than thought.
This doesn’t mean engineering mindset is better than scientific mindset, but it does mean that engineering mindset isn’t as bad as you say.
And we should also admit that the research program you propose (agent foundations) has been underwhelming at best, and even though problems like the genie knows, but does not care has unfortunately been more accurate than we thought (mostly because we scaled up RL again, and are likely to keep doing this because of incentives), solutions from agent foundations have been lacking.
IMO, the most exciting safety work (so far) is modelled on mundane solutions to exotic problems, like Risk-Averse AIs (with relevant experimental results here)
So in conclusion, while I could see engineering mindset being worse than scientific mindset, I also think that this is less obvious of a conclusion than you do (assuming the claim of engineering mindset being worse than scientific mindset is correct), and I wanted to present a partial positive case of engineering mindset to offset potential biases.
Science is partly a coupled process between engineering and theory, it is also hard to do causal inference from what counts as theory and what counts as practice.
Would you say that the discovery of the higgs boson was something that happened as a consequence through theoretical or practical physics?
What about deception, inner misalignment, general interpretability methods, if you trace their intelluectual lineage where do they come from? They’re not fully agent foundations but if you compare if they’re more ML based or coming from the larger space of theoretical alignment research I would attribute more causal influence to theoretical AI Safety. This is what agent foundations was up to like 3 years ago!
Or what do we mean by agent foundations here? What would you draw the boundaries around? Is it maybe better to use the word Theoretical AI Safety research?
Given specific assumptions about scientific progress where we can iteratively improve it and it is clear how we would even aim for it in the first place.
How much is this field biology and how much is it computer science? If it is computer science then it is bloody complicated combinatorial optimisation and dynamic programming. This field is in it’s philosophical nature closer to studying growing systems not engineered systems.
I would have much less beef with the engineering mindset if it hadn’t accelerated capabilities so much, as I’ll explain in the next post (and also if it weren’t so related to conceptually confused ML research, as I’ll explain in the post after that).
My main issue is The Counterfactual Quiet AGI Timeline. Scaling laws of neural nets had capabilities become more like treasures waiting for the right amount of compute. Once anyone invested the compute, the treasure would be his and everyone would rush for similar treasures, until one of them summons demons...
I don’t yet have a situation-global opinion here, but “once anyone invested the compute, the treasure would be theirs” doesn’t seem like a valid transition. The primary constraint in connecting abstract research to economic applications is generalist-capable domain-experts being motivated to bridge the gap from theory to something a VC or customer can understand.
While you can’t stay still forever relying on the absence of that force, a small non-prestigious community uniquely positioned to work in a field deciding not to seek investments and capability-research would certainly have further delayed the viability of the treasure, in turn supporting non-capabilities-dependent alignment-contribution tactics like intelligence augmentation, foundational research, philosophical research, or prestige-building from non-X-risk charitable efforts to acquire credibility for X-risk-mitigation.
There are valid replacement-reasons that could slot in the same place in your presented reasoning, including-but-not-limited-to longer timeline-estimates making personal-judgement-of-influencers more important than timeline-reductions, reason to believe some specific other party was already working on making transformer-based autonomously-operating AI economically viable at the time, expectation that capabilities could meaningfully help with other X-risks before they became a concern themselves, etc. I’m not sure which specifically of those you would agree with, though.