Isn’t the stronger counterargument orthogonality? Making models more capable of arbitrary goals seems highly unlikely to lead to safety unless you also steer them better, and if you believe that, all this hair-splitting is unnecessary.
Isn’t the stronger counterargument orthogonality? Making models more capable of arbitrary goals seems highly unlikely to lead to safety unless you also steer them better, and if you believe that, all this hair-splitting is unnecessary.