In defense of advancing AI in philosophy: if moral realism is true and our best guess of what’s Good is wrong, then alignment to that will be really bad, possibly worse than extinction.
So:
Try to align to our best guess: very low p(Good), medium x-risk, high p(survival | non-Good)
Make AIs figure it out: higher p(Good), high x-risk, medium p(survival | non-Good)
If getting the Good is all that matters, then option 2 beats 1 because the timelines were we don’t find the Good are not… good, regardless of survival.
In defense of advancing AI in philosophy: if moral realism is true and our best guess of what’s Good is wrong, then alignment to that will be really bad, possibly worse than extinction.
So:
Try to align to our best guess: very low p(Good), medium x-risk, high p(survival | non-Good)
Make AIs figure it out: higher p(Good), high x-risk, medium p(survival | non-Good)
If getting the Good is all that matters, then option 2 beats 1 because the timelines were we don’t find the Good are not… good, regardless of survival.