Suppose your judges were all random coinflips. That wouldn’t work to drive correct outputs for debate as a means of superalignment. Can you tell me what property the judges need, which random coinflips lack, and which isn’t ‘the judgments are statistically unbiased’, in order for this scheme to work?
Well, a fairly obvious ‘necessary property’ is ‘the judgments are correlated with correctness’ (that is, they provide Shannon information about whether outputs are correct or not). That’s a property that random coinflips don’t have, and nor do “judges who’ve never seen the Emperor”. (Of course, in the presence of bias, that might not be a sufficient property.)
It is not necessary, in general, for “the plurality or the majority [to] always [be] right” to extract accurate information out of judges that are sometimes wrong as individuals. After all, financial markets do just that, because any predictable pattern of wrongness becomes exploitable, and exploiting patterned noise tends to suppress that noise, increasing SNR (as long as there is any signal at all to focus). Whether this is in any way applicable to the ‘humans judging an Alignment debate between AIs’ case is unclear at best, but it does seem to show that there is not a general Theorem that accurate judgment-aggregation systems require the properties over their judges that you imply are required.
So I would have to say that you have, at least from the perspective of this reader, failed to make comprehensible just what Cat-Belling Problem you think Irving hasn’t addressed. (Perhaps it would help if you set the problem up formally so that your Impossibility Theorem can be stated mathematically?)
FTR I have no dog in this fight, no strong opinions either way about the ‘human-judged AI-debated Alignment plans’ plan.
Revisiting this nine years later, it occurs to me that that last equation is a corollary of the Law of Conservation of Expected Evidence; if you know how much a positive and a negative result from experiment B would shift your probability of A, then the probability of a positive result from experiment B must be such as to make the two arms balance. And that is precisely the expression.