I think controlling AI x-risk requires insight into the behavior of complex mathematical objects. To be a little less fuzzy, we need a working theory of mind. It is our mathematical competence which will determine how much control we have here.
I think RSI (a.k.a. self-improving AI autoresearchers) is inevitable for a wide variety of commercial, military, and/or geopolitical reasons, no matter what treaties are signed, no matter what promises are made. At least the NSA will do it. This is a curse in that it speeds the timeline. This is a blessing in that it gives us a chance. Again, though, the key to bounding the alignment delta between versions is going to be a theory of mind and its associated complex mathematics. If the delta is unbounded or even just too large, we lose and die. If the delta is small, we have a chance.
I think just waiting, or rather just researching, allows successive model versions to drift apart in mind space and makes it more difficult to bound the alignment delta. An all round catastrophe. We need to cognitively enhance each mathematician we have and put them to work developing (the mathematics behind) a theory of mind and/or on their favorite alignment subproblem.
I think controlling AI x-risk requires insight into the behavior of complex mathematical objects. To be a little less fuzzy, we need a working theory of mind. It is our mathematical competence which will determine how much control we have here.
I think RSI (a.k.a. self-improving AI autoresearchers) is inevitable for a wide variety of commercial, military, and/or geopolitical reasons, no matter what treaties are signed, no matter what promises are made. At least the NSA will do it. This is a curse in that it speeds the timeline. This is a blessing in that it gives us a chance. Again, though, the key to bounding the alignment delta between versions is going to be a theory of mind and its associated complex mathematics. If the delta is unbounded or even just too large, we lose and die. If the delta is small, we have a chance.
I think just waiting, or rather just researching, allows successive model versions to drift apart in mind space and makes it more difficult to bound the alignment delta. An all round catastrophe. We need to cognitively enhance each mathematician we have and put them to work developing (the mathematics behind) a theory of mind and/or on their favorite alignment subproblem.