You would never learn “The model is not trying to play chess well” (the word “trying” will never appear in any outputs from this system). But when you ask “How likely is the model to win this game of chess?”, you’ll get a lower number in games where the model is not trying (because the explanation necessarily captures the mechanisms by which “trying” causes the probability of victory to be higher).
Similarly for takeover attempts, you won’t learn “the model is trying to take over.” But if you ask your estimator “What’s the chance that something would look horribly wrong if I investigated this incident?” the probability will spike when the model starts trying to take over.
(Unless it takes over in such a way that you would not notice even if you investigated. To get around that you could try to say “Well what if I shut down all the AI and did a big investigation at the moment when it would be most damning?” I think that getting answers to this kind of question would be a huge deal and would get us most of what we want / would require solving the core difficulties in alignment. but it’s still far from bulletproof and at some point we need to roll up our sleeves and deal with the philosophical work I mentioned about referring to the causes for patterns in our observations and not just the observations themselves.)
You would never learn “The model is not trying to play chess well” (the word “trying” will never appear in any outputs from this system). But when you ask “How likely is the model to win this game of chess?”, you’ll get a lower number in games where the model is not trying (because the explanation necessarily captures the mechanisms by which “trying” causes the probability of victory to be higher).
Similarly for takeover attempts, you won’t learn “the model is trying to take over.” But if you ask your estimator “What’s the chance that something would look horribly wrong if I investigated this incident?” the probability will spike when the model starts trying to take over.
(Unless it takes over in such a way that you would not notice even if you investigated. To get around that you could try to say “Well what if I shut down all the AI and did a big investigation at the moment when it would be most damning?” I think that getting answers to this kind of question would be a huge deal and would get us most of what we want / would require solving the core difficulties in alignment. but it’s still far from bulletproof and at some point we need to roll up our sleeves and deal with the philosophical work I mentioned about referring to the causes for patterns in our observations and not just the observations themselves.)