Training on alignment evals as opposed to capabilities-eliciting tasks is one of the worst things a lab could do since it means a loss of a way to test alignment similarly to how the Most Forbidden Technique teaches models to hide misalignment from the CoT. On the other hand, any lab, even the one which solved alignment or believed that it solved alignment, would train the models on capabilities-eliciting tasks.
You linked to Ngo’s post. However, suppose that you have a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine” and DeepCent has the same kill machine with a 50% chance of commiting genocide. Then we would have to rule out DeepCent activating its machine. I described similar issues in my post.
I do not think that 1. is very relevant to the central point. Maybe that was a bad example. However, even if the lab isn’t directly RLing on that alignment eval or whatever, they may be “grad student descent”-ing up the eval and achieving a similar effect. Either way, I think the result is a increase by X% of entering an aligned RSI flywheel and a decrease by Y% of everyone freaking out and slowing down.
I think 2. is just the coordination problem I describe? I did not read your post, so I don’t know if you are pointing to something other than the fact that being noble in order to encourage coordination on this would have to account for the fact that this coordination must extend to China.
This conflates two issues:
Training on alignment evals as opposed to capabilities-eliciting tasks is one of the worst things a lab could do since it means a loss of a way to test alignment similarly to how the Most Forbidden Technique teaches models to hide misalignment from the CoT. On the other hand, any lab, even the one which solved alignment or believed that it solved alignment, would train the models on capabilities-eliciting tasks.
You linked to Ngo’s post. However, suppose that you have a community that is saying, “well, 20% chance the kill machine kills us all but it would be 35% if we weren’t taping up the wiring in the kill machine” and DeepCent has the same kill machine with a 50% chance of commiting genocide. Then we would have to rule out DeepCent activating its machine. I described similar issues in my post.
I do not think that 1. is very relevant to the central point. Maybe that was a bad example. However, even if the lab isn’t directly RLing on that alignment eval or whatever, they may be “grad student descent”-ing up the eval and achieving a similar effect. Either way, I think the result is a increase by X% of entering an aligned RSI flywheel and a decrease by Y% of everyone freaking out and slowing down.
I think 2. is just the coordination problem I describe? I did not read your post, so I don’t know if you are pointing to something other than the fact that being noble in order to encourage coordination on this would have to account for the fact that this coordination must extend to China.