That’s an interesting question. I think my answer would be, from a perspective that can’t directly appreciate (and doubts!) the benefit of a lot of abstract theory, I can still recognize that it’s low cost (theorists are cheap) and worth investing in, because theory in general is valuable even if I personally have quite an imperfect sense of which theory will turn out to bear fruit. I do, for instance, obviously see value in philosophy and pure mathematics!
And I even think I have some ability to discern what theory to prioritize, looking for things like goal-relevance, substance, creativity, beauty, technical prowess, etc. (Though of course I also have to rely on expert advice.)
Where it gets gnarly is policy: theoretical considerations for costly policies are often going to strike me as asking people to pay real costs for dubious benefits. The different perspectives have fundamentally different epistemological methods and I’m uncertain where I stand. That’s why I don’t work in policy and why, despite being candid about my perspective sometimes, I try not to make authoritative-sounding pronouncements about what should be done. I really don’t want to cause harm by being wrong, and at times I regret that I can’t always resist shooting my mouth off.
I also think it makes sense to aim for Pareto improvements—prioritizing things that reduce risk a lot relative to their cost. Fundamental safety-relevant theoretical research is one of them. “Mundane” security & resilience work is another. Developing standards for AI safety that can be operationalized as requirements, which IMO will require “deconfusion” and new conceptual vocabulary, seems good to me. I’m also interested in thinking about how we could get the economic & scientific & even social/institutional benefits of actually-existing AI at lower risk, which I haven’t heard almost anyone discuss. (Though IFP just came out with a discussion of “pacing” strategies that seems to be engaging with this, for instance.)
It does seem like, by definition, policy has to be built out of “legible” pieces.
(Well, the main alternative is “somehow, someone with great taste ends up in charge.” But, the people who don’t share that taste would be pretty sus. My story for “illegible taste decisionmaking” working out would involve people with very “scoped power”, who somehow are entrusted to make calls about what AI is safe to scale, that is structured such that nobody has to worry about them wielding that power to do other bad things).
But, assuming for now “we don’t get illegibly good decisionmakers with power over AI training”...
...one thing that comes immediately to mind is “I think almost nobody (modulo a few extropians) actually wants hard/fast takeoff.” People who don’t want AI regulation mostly aren’t expecting hard takeoff in the first place, and are assuming the natural low-regulated course of things would be a technology that’s manageable.
AI 2040 paints a story where takeoff happens over 10 years and it’s still pretty crazy feeling by any reasonable standard. Many people were alarmed that the “doomers” expected faster takeoff than the accelerationists even when dramatically slowing down.
This leaves the problem of disagreement over “do we need to stop research, because we’re rapidly approaching a critical threshold where you get godlike AI that we aren’t prepared for?”. But, if there ends up being any kind warning shot or clarity that, yeah actually godlike AI is right on the horizon, I kinda just expect a lot of agreement?
(I don’t actually remember: are you like still unsure if you expect godlike AI at all, or expecting it like multiple years after humanlevel conceptual research? I’m not actually sure which parts feel fake to you atm)
A lot of people I respect don’t believe in “godlike” AI at all, so I am confused there.
I also worry about the policy implementation of a “pause” being done in the wrong way with lots of side effects.
I agree that if we have the “slow” AI progress laid out in AI 2040 that’s actually way faster than late-20c/early-21c tech progress has been, then I don’t have anything to complain about from an economic growth point of view.
Frankly it’s weird to me that the labs chose to go to RSI as fast as possible in the first place; intuitively I would have thought the natural thing to do would be to take your time and get it right and get a much better understanding of how neural nets even work.
Frankly it’s weird to me that the labs chose to go to RSI as fast as possible in the first place; intuitively I would have thought the natural thing to do would be to take your time and get it right and get a much better understanding of how neural nets even work.
Really? What do you find so confusing about it?
It’s the kind of thing you do if you want to affect the world through power. It’s quite uncertain how far we are from understanding NNs well (anywhere from 1-10 years seems reasonable, even with our current AI helpers). And in the meantime someone else could invest in scaling NNs and get there, or you won’t be able to sustain enough expectation of future revenue for your investment multiples, or if you’re OpenAI then Anthropic will do it first and you lose power. (Though people are spooked enough by recent incidents that these two labs might start pacing the frontier now, so maybe your intuitions are more correct.)
That’s an interesting question. I think my answer would be, from a perspective that can’t directly appreciate (and doubts!) the benefit of a lot of abstract theory, I can still recognize that it’s low cost (theorists are cheap) and worth investing in, because theory in general is valuable even if I personally have quite an imperfect sense of which theory will turn out to bear fruit. I do, for instance, obviously see value in philosophy and pure mathematics!
And I even think I have some ability to discern what theory to prioritize, looking for things like goal-relevance, substance, creativity, beauty, technical prowess, etc. (Though of course I also have to rely on expert advice.)
Where it gets gnarly is policy: theoretical considerations for costly policies are often going to strike me as asking people to pay real costs for dubious benefits. The different perspectives have fundamentally different epistemological methods and I’m uncertain where I stand. That’s why I don’t work in policy and why, despite being candid about my perspective sometimes, I try not to make authoritative-sounding pronouncements about what should be done. I really don’t want to cause harm by being wrong, and at times I regret that I can’t always resist shooting my mouth off.
I also think it makes sense to aim for Pareto improvements—prioritizing things that reduce risk a lot relative to their cost. Fundamental safety-relevant theoretical research is one of them. “Mundane” security & resilience work is another. Developing standards for AI safety that can be operationalized as requirements, which IMO will require “deconfusion” and new conceptual vocabulary, seems good to me. I’m also interested in thinking about how we could get the economic & scientific & even social/institutional benefits of actually-existing AI at lower risk, which I haven’t heard almost anyone discuss. (Though IFP just came out with a discussion of “pacing” strategies that seems to be engaging with this, for instance.)
Cool.
It does seem like, by definition, policy has to be built out of “legible” pieces.
(Well, the main alternative is “somehow, someone with great taste ends up in charge.” But, the people who don’t share that taste would be pretty sus. My story for “illegible taste decisionmaking” working out would involve people with very “scoped power”, who somehow are entrusted to make calls about what AI is safe to scale, that is structured such that nobody has to worry about them wielding that power to do other bad things).
But, assuming for now “we don’t get illegibly good decisionmakers with power over AI training”...
...one thing that comes immediately to mind is “I think almost nobody (modulo a few extropians) actually wants hard/fast takeoff.” People who don’t want AI regulation mostly aren’t expecting hard takeoff in the first place, and are assuming the natural low-regulated course of things would be a technology that’s manageable.
AI 2040 paints a story where takeoff happens over 10 years and it’s still pretty crazy feeling by any reasonable standard. Many people were alarmed that the “doomers” expected faster takeoff than the accelerationists even when dramatically slowing down.
This leaves the problem of disagreement over “do we need to stop research, because we’re rapidly approaching a critical threshold where you get godlike AI that we aren’t prepared for?”. But, if there ends up being any kind warning shot or clarity that, yeah actually godlike AI is right on the horizon, I kinda just expect a lot of agreement?
(I don’t actually remember: are you like still unsure if you expect godlike AI at all, or expecting it like multiple years after humanlevel conceptual research? I’m not actually sure which parts feel fake to you atm)
A lot of people I respect don’t believe in “godlike” AI at all, so I am confused there.
I also worry about the policy implementation of a “pause” being done in the wrong way with lots of side effects.
I agree that if we have the “slow” AI progress laid out in AI 2040 that’s actually way faster than late-20c/early-21c tech progress has been, then I don’t have anything to complain about from an economic growth point of view.
Frankly it’s weird to me that the labs chose to go to RSI as fast as possible in the first place; intuitively I would have thought the natural thing to do would be to take your time and get it right and get a much better understanding of how neural nets even work.
Really? What do you find so confusing about it?
It’s the kind of thing you do if you want to affect the world through power. It’s quite uncertain how far we are from understanding NNs well (anywhere from 1-10 years seems reasonable, even with our current AI helpers). And in the meantime someone else could invest in scaling NNs and get there, or you won’t be able to sustain enough expectation of future revenue for your investment multiples, or if you’re OpenAI then Anthropic will do it first and you lose power. (Though people are spooked enough by recent incidents that these two labs might start pacing the frontier now, so maybe your intuitions are more correct.)
Can you name 5? Who also have clearly thought about the topic in a serious way? (E.g. not Tyler Cowen)
Or is the issue with “godlike”, but they’d be fine with “wildly superhuman”.