Regarding question 1: “Uncooperativeness” is sort-of a free parameter here, at least in the way the model functions. If you can build an evaluation whose quality scales sufficiently quickly with AI capability, you can use that as your effective measure of uncooperativeness. I think for the policy feedback aim, credibility and consensus are also really important. So the practical limitation is more about what kind of evaluation pipelines you can build, and what kind of consensus you can build around them.
I am personally sympathetic to more ambitious “alignment” targets, but I think consensus is may be a limiting factor on how ambitious you can get.
That said, it’s not entirely a free parameter; if you don’t take a stand on traits that are effectively self propagating then you may add a bunch of uncertainty to the self propagation term, which makes the model worse. Similarly if you know a trait undermines your monitoring efforts then not calling it “uncooperative” basically invalidates the broader project.
Regarding question 2: Yes, I’m assuming the detector is not becoming less precise (or becoming more precise); I breezily swept that under “regularity conditions”. But when I think about answering your question, I wonder if it wouldn’t be better to get a sequence of judgement diffs between older and newer detectors, and then you could track the “true positive according to SOTA” and “false positive according to SOTA” trends separately.
Regarding question 1: “Uncooperativeness” is sort-of a free parameter here, at least in the way the model functions. If you can build an evaluation whose quality scales sufficiently quickly with AI capability, you can use that as your effective measure of uncooperativeness. I think for the policy feedback aim, credibility and consensus are also really important. So the practical limitation is more about what kind of evaluation pipelines you can build, and what kind of consensus you can build around them.
I am personally sympathetic to more ambitious “alignment” targets, but I think consensus is may be a limiting factor on how ambitious you can get.
That said, it’s not entirely a free parameter; if you don’t take a stand on traits that are effectively self propagating then you may add a bunch of uncertainty to the self propagation term, which makes the model worse. Similarly if you know a trait undermines your monitoring efforts then not calling it “uncooperative” basically invalidates the broader project.
Regarding question 2: Yes, I’m assuming the detector is not becoming less precise (or becoming more precise); I breezily swept that under “regularity conditions”. But when I think about answering your question, I wonder if it wouldn’t be better to get a sequence of judgement diffs between older and newer detectors, and then you could track the “true positive according to SOTA” and “false positive according to SOTA” trends separately.