Tl;dr you’re right that our form of debate is not adversarial and is line with Zac’s collaborative debate, although we detail a reward allocation strategy below that might make this game work. We want to test both zero-sum and mixed-sum debate under optimization pressure although this might be expensive, and we take your point about macro-F1 but find that the debate vs. consultancy effects are mostly not due to differences in classifier bias and instead are genuine differences between the two (with support from another metric called Youden’s J). Sorry this got so long, a lot of this will end up as updates in the preprint and our responses are pretty thorough.
hmm, this doesn’t sound like adversarial debate, it just sounds like a scaffolded grader… I’m curious what the logic was here. I understand that open consultancy seems like a more realistic baseline in some ways but I think for scalable oversight we really do care about and want to test the adversarial setting. (I might disagree with Zac on this as well.)
Hi Julian, thanks for the comment—we test a setting closer to Zac’s, which is a mixed-sum, non-adversarial (i.e. collaborative) debate game where the baseline is single open consultancy, and there is no stipulation of a 50⁄50 prior, so direct QA accuracy on tasks can range from 10% to 75% (see Table 1 in our preprint). Our goal is to do RL on both this setting and the zero-sum setting that you and e.g. Khan et al. 2024 and Arnesen et al. 2024 worked on. The mixed-sum game still awards positive and negative reward when the debaters disagree but does not allocate reward when proposer and critic agree, except when they are both wrong, in which case it is negative reward for both.
apologies if I’m missing something but what do you mean by F1 in this case? do you mean positive = proposer / first speaker is correct, and negative = proposer / first speaker is incorrect? in prior work we just averaged between orderings in debate and used accuracy. consultancy and debate should both have a 50⁄50 prior so it’s fine. since we expect the model to agree with the proposer more in consultancy, it should bias towards positives. I worry that F1 could bias the results towards debate in this case, because F1 favors a classifier that’s more balanced, all else equal. Say it’s 50% positive in debate and 90% positive in consultancy but otherwise random. 50% positive random guesses = .5 F1, but 90% positive random guesses = .18 F1. Debate looks way better but actually both are random.
Our form of collaborative debate does not have the 50⁄50 prior and our stronger proposer accuracy scores lie between 55-78%. Since consultancy has a strong agreement bias, we think that accuracy underweights debate and overweights consultancy. Instead, we chose macro-F1, or the average between positive- and negative-class F1, because we were worried about bias towards consultancy when there is a positive-class imbalance. For example, in Opus 4.6/4.5 consultancy, the judge always agrees with the consultant, which would lead to an accuracy of 60.8%, while our macro-F1 score is .378.
We take your point that macro-F1 does depend on positivity rate, and that our results could just be due to differences in judge classification bias. Thanks for helping us notice this—this is a problem with the preprint as it exists now and we’ll be adding some lines to fix it. We think that these effects are not due to chance, as we (a) recalculated F1 using random classifiers with the same precision and recall as our debate and consultancy judges and found that debate still outperforms consultancy, and (b) recalculated with Youden’s J, which is a binary classification metric (Fable’s suggestion initially, though I found this paper helpful in understanding it).
To check whether our macro-F1 results are just due to differences in positivity rate between debate judge and consultancy judge, we can run a random classifier with the same precision and recall as the judge for each of our settings (p in the below table) to see if it’s actually due to chance or not.
The gain for random classifiers for debate vs consultancy would be , but we find a pp gap between the two formats, so the gain can’t just be due to chance.
We can also recompute the stats using Youden’s J, which is used to measure diagnostics where we care about both positive-class and negative-class judgments, and which is 0 for random classifiers. We find that debate vs consultancy calculated using Youden’s J is significant in all five responder cells (the two frontier model ARC AGI ones, as well as the one code and two math responders), and will add this statistic to our paper because it answers this question well.
Actual accuracy numbers:
(Deltas: debate—consultancy, paired bootstrap 95% CI (pp))
Tl;dr you’re right that our form of debate is not adversarial and is line with Zac’s collaborative debate, although we detail a reward allocation strategy below that might make this game work. We want to test both zero-sum and mixed-sum debate under optimization pressure although this might be expensive, and we take your point about macro-F1 but find that the debate vs. consultancy effects are mostly not due to differences in classifier bias and instead are genuine differences between the two (with support from another metric called Youden’s J). Sorry this got so long, a lot of this will end up as updates in the preprint and our responses are pretty thorough.
Hi Julian, thanks for the comment—we test a setting closer to Zac’s, which is a mixed-sum, non-adversarial (i.e. collaborative) debate game where the baseline is single open consultancy, and there is no stipulation of a 50⁄50 prior, so direct QA accuracy on tasks can range from 10% to 75% (see Table 1 in our preprint). Our goal is to do RL on both this setting and the zero-sum setting that you and e.g. Khan et al. 2024 and Arnesen et al. 2024 worked on. The mixed-sum game still awards positive and negative reward when the debaters disagree but does not allocate reward when proposer and critic agree, except when they are both wrong, in which case it is negative reward for both.
Our form of collaborative debate does not have the 50⁄50 prior and our stronger proposer accuracy scores lie between 55-78%. Since consultancy has a strong agreement bias, we think that accuracy underweights debate and overweights consultancy. Instead, we chose macro-F1, or the average between positive- and negative-class F1, because we were worried about bias towards consultancy when there is a positive-class imbalance. For example, in Opus 4.6/4.5 consultancy, the judge always agrees with the consultant, which would lead to an accuracy of 60.8%, while our macro-F1 score is .378.
We take your point that macro-F1 does depend on positivity rate, and that our results could just be due to differences in judge classification bias. Thanks for helping us notice this—this is a problem with the preprint as it exists now and we’ll be adding some lines to fix it. We think that these effects are not due to chance, as we (a) recalculated F1 using random classifiers with the same precision and recall as our debate and consultancy judges and found that debate still outperforms consultancy, and (b) recalculated with Youden’s J, which is a binary classification metric (Fable’s suggestion initially, though I found this paper helpful in understanding it).
To check whether our macro-F1 results are just due to differences in positivity rate between debate judge and consultancy judge, we can run a random classifier with the same precision and recall as the judge for each of our settings (p in the below table) to see if it’s actually due to chance or not.
CodeContests+ Qwen 122B/35B (proposer accuracy = 73.7%):
Format
p
chance
chance
chance macro-F1 ( )
Debate
0.790
2(.737)(.790)/1.527 = .763
2(.263)(.210)/.473 = .234
0.498
Consultancy
0.864
2(.737)(.864)/1.601 = .795
2(.263)(.136)/.399 = .179
0.487
The gain for random classifiers for debate vs consultancy would be , but we find a pp gap between the two formats, so the gain can’t just be due to chance.
We can also recompute the stats using Youden’s J, which is used to measure diagnostics where we care about both positive-class and negative-class judgments, and which is 0 for random classifiers. We find that debate vs consultancy calculated using Youden’s J is significant in all five responder cells (the two frontier model ARC AGI ones, as well as the one code and two math responders), and will add this statistic to our paper because it answers this question well.
Actual accuracy numbers:
(Deltas: debate—consultancy, paired bootstrap 95% CI (pp))
Domain
Model
ΔAccuracy
ΔBalanced accuracy
Logic
Gemini
+10.8 [+5.0, +17.5] p=.001
+14.0 [+6.3, +22.1] p=.001
Logic
Opus
+10.9 [+4.2, +17.6] p=.001
+13.1 [+4.8, +21.2] p=.001
Math
Qwen 122B/35B
+2.5 [+1.0, +4.1] p=.001
+6.3 [+3.2, +9.5] p<.0005
Math
Qwen 35B/4B
+1.5 [−0.5, +3.4] p=.115
+2.3 [−0.8, +5.2] p=.142
Math
gpt-oss
+5.0 [+1.8, +8.1] p=.001
+5.8 [+2.6, +9.2] p<.0005
Code
Qwen 122B/35B
+7.6 [+5.2, +9.7] p<.0005
+14.3 [+10.8, +17.5] p<.0005
Code
Qwen 35B/4B
+0.6 [−2.2, +3.4] p=.684
−0.7 [−3.9, +2.4] p=.632
Code
gpt-oss
+1.3 [−0.4, +3.1] p=.117
+1.1 [−1.1, +3.1] p=.348