I’m curious how exactly the “if a model simply decides not to play chess well in certain situations and you understand how the model works, you should be able to recognize when that happens” would work.
In traditional interability, a human can just read it the goals. With mechanistic explanations, is it that you could show that fiddling with certain parameters or activations would increase the win-rate (without needing to sample it)? Or that if it has other goals you could find an explanation of why it is following those other goals, thus proving it isn’t always aligned to chess?
You would never learn “The model is not trying to play chess well” (the word “trying” will never appear in any outputs from this system). But when you ask “How likely is the model to win this game of chess?”, you’ll get a lower number in games where the model is not trying (because the explanation necessarily captures the mechanisms by which “trying” causes the probability of victory to be higher).
Similarly for takeover attempts, you won’t learn “the model is trying to take over.” But if you ask your estimator “What’s the chance that something would look horribly wrong if I investigated this incident?” the probability will spike when the model starts trying to take over.
(Unless it takes over in such a way that you would not notice even if you investigated. To get around that you could try to say “Well what if I shut down all the AI and did a big investigation at the moment when it would be most damning?” I think that getting answers to this kind of question would be a huge deal and would get us most of what we want / would require solving the core difficulties in alignment. but it’s still far from bulletproof and at some point we need to roll up our sleeves and deal with the philosophical work I mentioned about referring to the causes for patterns in our observations and not just the observations themselves.)
Interesting!
I’m curious how exactly the “if a model simply decides not to play chess well in certain situations and you understand how the model works, you should be able to recognize when that happens” would work.
In traditional interability, a human can just read it the goals. With mechanistic explanations, is it that you could show that fiddling with certain parameters or activations would increase the win-rate (without needing to sample it)? Or that if it has other goals you could find an explanation of why it is following those other goals, thus proving it isn’t always aligned to chess?
You would never learn “The model is not trying to play chess well” (the word “trying” will never appear in any outputs from this system). But when you ask “How likely is the model to win this game of chess?”, you’ll get a lower number in games where the model is not trying (because the explanation necessarily captures the mechanisms by which “trying” causes the probability of victory to be higher).
Similarly for takeover attempts, you won’t learn “the model is trying to take over.” But if you ask your estimator “What’s the chance that something would look horribly wrong if I investigated this incident?” the probability will spike when the model starts trying to take over.
(Unless it takes over in such a way that you would not notice even if you investigated. To get around that you could try to say “Well what if I shut down all the AI and did a big investigation at the moment when it would be most damning?” I think that getting answers to this kind of question would be a huge deal and would get us most of what we want / would require solving the core difficulties in alignment. but it’s still far from bulletproof and at some point we need to roll up our sleeves and deal with the philosophical work I mentioned about referring to the causes for patterns in our observations and not just the observations themselves.)