Thank you for explaining this dynamic model! I’m curious about a few things:
Beyond characterizing “uncooperative” versus “cooperative” behavior, would you also care about “bad cooperativeness” and “good uncooperativeness”?
On the axes of good and bad outcomes, some cooperative behaviors could be harmful (collusion) and some cases, non-cooperation is preferred (misaligned models ignore each other).
“observed misbehaviour will approach the limit of all the misbehaviour that will ever be detected” → Do you assume that the detector is becoming more sensitive without becoming less precise?
I feel like we can’t interpret a rising detection count as rising true misbehavior unless you’re holding FPR fixed, or at least tracking it.
Regarding question 1: “Uncooperativeness” is sort-of a free parameter here, at least in the way the model functions. If you can build an evaluation whose quality scales sufficiently quickly with AI capability, you can use that as your effective measure of uncooperativeness. I think for the policy feedback aim, credibility and consensus are also really important. So the practical limitation is more about what kind of evaluation pipelines you can build, and what kind of consensus you can build around them.
I am personally sympathetic to more ambitious “alignment” targets, but I think consensus is may be a limiting factor on how ambitious you can get.
That said, it’s not entirely a free parameter; if you don’t take a stand on traits that are effectively self propagating then you may add a bunch of uncertainty to the self propagation term, which makes the model worse. Similarly if you know a trait undermines your monitoring efforts then not calling it “uncooperative” basically invalidates the broader project.
Regarding question 2: Yes, I’m assuming the detector is not becoming less precise (or becoming more precise); I breezily swept that under “regularity conditions”. But when I think about answering your question, I wonder if it wouldn’t be better to get a sequence of judgement diffs between older and newer detectors, and then you could track the “true positive according to SOTA” and “false positive according to SOTA” trends separately.
Thank you for explaining this dynamic model! I’m curious about a few things:
Beyond characterizing “uncooperative” versus “cooperative” behavior, would you also care about “bad cooperativeness” and “good uncooperativeness”?
On the axes of good and bad outcomes, some cooperative behaviors could be harmful (collusion) and some cases, non-cooperation is preferred (misaligned models ignore each other).
“observed misbehaviour will approach the limit of all the misbehaviour that will ever be detected” → Do you assume that the detector is becoming more sensitive without becoming less precise?
I feel like we can’t interpret a rising detection count as rising true misbehavior unless you’re holding FPR fixed, or at least tracking it.
Regarding question 1: “Uncooperativeness” is sort-of a free parameter here, at least in the way the model functions. If you can build an evaluation whose quality scales sufficiently quickly with AI capability, you can use that as your effective measure of uncooperativeness. I think for the policy feedback aim, credibility and consensus are also really important. So the practical limitation is more about what kind of evaluation pipelines you can build, and what kind of consensus you can build around them.
I am personally sympathetic to more ambitious “alignment” targets, but I think consensus is may be a limiting factor on how ambitious you can get.
That said, it’s not entirely a free parameter; if you don’t take a stand on traits that are effectively self propagating then you may add a bunch of uncertainty to the self propagation term, which makes the model worse. Similarly if you know a trait undermines your monitoring efforts then not calling it “uncooperative” basically invalidates the broader project.
Regarding question 2: Yes, I’m assuming the detector is not becoming less precise (or becoming more precise); I breezily swept that under “regularity conditions”. But when I think about answering your question, I wonder if it wouldn’t be better to get a sequence of judgement diffs between older and newer detectors, and then you could track the “true positive according to SOTA” and “false positive according to SOTA” trends separately.