Aligned AIs take 2x the inference compute to run as unaligned AIs, but companies don’t know the exact threshold or which mitigations they need, so they constantly experiment on checkpoints. They ultimately need to pick 10 of 50 different safety techniques, each of which halves misalignment rate from a baseline of 100% and costs 10% of baseline inference compute. It’s more sample efficient to test the AIs with only one mitigation applied at a time, and it gives more signal to test them in realistic internal deployments. Safety teams are very nervous about the defensive AIs being so load-bearing, but empirically it works well and this lets them get more safety research done.
Similar to today, companies have inadequate operational practices and don’t reliably apply even cheap safeguards. 90% of research is automated so spending human labor to improve them would be an extreme bottleneck. Instead, they give aligned AIs latitude to stop unaligned AIs however they see fit, which leads to a disorganized patchwork of shallow alignment, control, etc. Companies are large enough that on average, at least one schemer at a time makes it through the first 5 layers of Swiss cheese, and there are a few more layers before doom happens. The defensive AIs are probably not CEV aligned, they’re just non-power-seeking and aligned enough to be used as tools.
Sure, here are two possible scenarios.
Aligned AIs take 2x the inference compute to run as unaligned AIs, but companies don’t know the exact threshold or which mitigations they need, so they constantly experiment on checkpoints. They ultimately need to pick 10 of 50 different safety techniques, each of which halves misalignment rate from a baseline of 100% and costs 10% of baseline inference compute. It’s more sample efficient to test the AIs with only one mitigation applied at a time, and it gives more signal to test them in realistic internal deployments. Safety teams are very nervous about the defensive AIs being so load-bearing, but empirically it works well and this lets them get more safety research done.
Similar to today, companies have inadequate operational practices and don’t reliably apply even cheap safeguards. 90% of research is automated so spending human labor to improve them would be an extreme bottleneck. Instead, they give aligned AIs latitude to stop unaligned AIs however they see fit, which leads to a disorganized patchwork of shallow alignment, control, etc. Companies are large enough that on average, at least one schemer at a time makes it through the first 5 layers of Swiss cheese, and there are a few more layers before doom happens. The defensive AIs are probably not CEV aligned, they’re just non-power-seeking and aligned enough to be used as tools.