I’m curious whether the military / government would choose to report alignment failures in the case of small scale catastrophes.
For example, let’s say there’s an alignment failure in an AI targeting system that mis-targets and kills something like hundreds-to-thousands of people. Failing to acknowledge the mistake would make the military seem indifferent / evil to many people. Acknowledging the mistake makes them seem less evil, but also makes them seem incompetent. I’m curious which messaging strategy they would prefer.
I agree that not acknowledging it at all would be counterproductive for the military. But I worry about obfuscation of the kind we saw in the Iran school bombing. They could acknowledge the initial error but not its mechanism. If the mistargeting is an outcome of misalignment (either actual power-seeking or just overeagerness), I have very little confidence that it’d be surfaced as such (edit: under current transparency standards), and not some hodge-podge of collateral damage and “a person did actually see it, that person made an error, our oversight was just bad—oops” and maybe “an AI system went a bit far here, but we’ve worked with the vendor to fix it and it is all good now, nothing to see here”
I’m curious whether the military / government would choose to report alignment failures in the case of small scale catastrophes.
For example, let’s say there’s an alignment failure in an AI targeting system that mis-targets and kills something like hundreds-to-thousands of people. Failing to acknowledge the mistake would make the military seem indifferent / evil to many people. Acknowledging the mistake makes them seem less evil, but also makes them seem incompetent. I’m curious which messaging strategy they would prefer.
I agree that not acknowledging it at all would be counterproductive for the military. But I worry about obfuscation of the kind we saw in the Iran school bombing. They could acknowledge the initial error but not its mechanism. If the mistargeting is an outcome of misalignment (either actual power-seeking or just overeagerness), I have very little confidence that it’d be surfaced as such (edit: under current transparency standards), and not some hodge-podge of collateral damage and “a person did actually see it, that person made an error, our oversight was just bad—oops” and maybe “an AI system went a bit far here, but we’ve worked with the vendor to fix it and it is all good now, nothing to see here”