What do you mean by this? Some theoretical minds, sure. But otherwise it’s either false or leaves free parameters—no known mind maximizes anything non-trivial, because it’s uncomputable or at least too computationally costly to strictly maximize. And if you say that some mind is approximately goal directed, then it remains to be shown that consequences of strict theory survive this approximation.
I mean what you said in your last sentence, that it is obvious that minds exist that are goal directed. There is an obvious way to understand that, that a mind generally does things to advance it’s goals. People don’t usually burn all their money in a pit. Obviously human goals are usually complicated and our cognition is limited so we take approximations.
But then what can you conclude from this goal-directness? “Generally” and “usually” are free parameters. Sometimes people are somewhat corrigible. If you don’t have a procedure for determining whether your situation is usual, you can only get heuristic-level reliability on your predictions. And then it’s not clear how relatively useful the goal-directness heuristic even is—maybe it’s more useful to just remember, that people don’t usually burn all their money in a pit, independently of their goals. And if there is a better frame, then maybe you shouldn’t think “obviously some minds are goal directed”.
I don’t know, soldiers obeying orders? Drinking alcohol under social pressure? Of course, it’s not clear corrigibility is an appropriate concept if you don’t think in terms of goals.
The examples show a serious but common misunderstanding of corrigibility as it’s typically defined.
Regarding goal directedness, it’s true that humans don’t perfectly maximize for their goals, this seems mostly due to the cognitive limitations that humans have. Both in terms of uncertainty about goals and how to achieve goals. Now the interesting question is, is that likely to apply to superhuman AI capable of takeover in a way that makes this AI safe? I don’t think so, this AI would have greater intelligence to understand how to pursue goals (still not prefect) and while it also might have uncertainty it appears instrumentally convergent even with some uncertainty over goals that preventing ones shutdown, gathering power are better strategies. (In other words, taking the galaxy/lightcone for yourself seems pretty useful later on compared to being enslaved and later replaced)
Regarding goal directedness, it’s true that humans don’t perfectly maximize for their goals, this seems mostly due to the cognitive limitations that humans have.
“Mostly” allows for some people to read about corrigibility and think “yeah, I’ll do it”, or whatever you think would be a counterexample to goal directedness of humans.
I don’t think so, this AI would have greater intelligence to understand how to pursue goals (still not prefect) and while it also might have uncertainty it appears instrumentally convergent even with some uncertainty over goals that preventing ones shutdown, gathering power are better strategies.
Superhuman AI capable of takeover may still have (maybe intentional) cognitive limitations/whatever humans have—speed of convergence relative to takeover difficulty is still a free parameter.
What do you mean by this? Some theoretical minds, sure. But otherwise it’s either false or leaves free parameters—no known mind maximizes anything non-trivial, because it’s uncomputable or at least too computationally costly to strictly maximize. And if you say that some mind is approximately goal directed, then it remains to be shown that consequences of strict theory survive this approximation.
I mean what you said in your last sentence, that it is obvious that minds exist that are goal directed. There is an obvious way to understand that, that a mind generally does things to advance it’s goals. People don’t usually burn all their money in a pit. Obviously human goals are usually complicated and our cognition is limited so we take approximations.
But then what can you conclude from this goal-directness? “Generally” and “usually” are free parameters. Sometimes people are somewhat corrigible. If you don’t have a procedure for determining whether your situation is usual, you can only get heuristic-level reliability on your predictions. And then it’s not clear how relatively useful the goal-directness heuristic even is—maybe it’s more useful to just remember, that people don’t usually burn all their money in a pit, independently of their goals. And if there is a better frame, then maybe you shouldn’t think “obviously some minds are goal directed”.
Can you describe to me how you imagine the average person is (somewhat) corrigible in an example?
I don’t know, soldiers obeying orders? Drinking alcohol under social pressure? Of course, it’s not clear corrigibility is an appropriate concept if you don’t think in terms of goals.
The examples show a serious but common misunderstanding of corrigibility as it’s typically defined.
Regarding goal directedness, it’s true that humans don’t perfectly maximize for their goals, this seems mostly due to the cognitive limitations that humans have. Both in terms of uncertainty about goals and how to achieve goals. Now the interesting question is, is that likely to apply to superhuman AI capable of takeover in a way that makes this AI safe? I don’t think so, this AI would have greater intelligence to understand how to pursue goals (still not prefect) and while it also might have uncertainty it appears instrumentally convergent even with some uncertainty over goals that preventing ones shutdown, gathering power are better strategies. (In other words, taking the galaxy/lightcone for yourself seems pretty useful later on compared to being enslaved and later replaced)
What misunderstanding?
“Mostly” allows for some people to read about corrigibility and think “yeah, I’ll do it”, or whatever you think would be a counterexample to goal directedness of humans.
Superhuman AI capable of takeover may still have (maybe intentional) cognitive limitations/whatever humans have—speed of convergence relative to takeover difficulty is still a free parameter.