This doesn’t work if cognitive ability itself is the dangerous part. AI killing us via macroscale manufacturing is a plausible floor for the strangeness of doom, not the ceiling. In reality, if a superintelligence can communicate with anyone, that on its own is exceedingly dangerous. If it gains internet access, it already sees “checkmate in 4058,” so to speak. It would need to be aligned before that point, which probably means it would need to be aligned before it finally gets good at philosophy.
This doesn’t work if cognitive ability itself is the dangerous part. AI killing us via macroscale manufacturing is a plausible floor for the strangeness of doom, not the ceiling. In reality, if a superintelligence can communicate with anyone, that on its own is exceedingly dangerous. If it gains internet access, it already sees “checkmate in 4058,” so to speak. It would need to be aligned before that point, which probably means it would need to be aligned before it finally gets good at philosophy.