Thanks for the comment! I looked into this and found that the length of the prompt does sometimes amplify the bias, but it doesn’t seem to be the cause of it. For the condition with the largest length spread (Sonnet on scheming), longer comparisons actually showed less bias (+49pp --> +37pp). Further, I’d expect this to be a relatively limited effect given that our longest prompts only reached around 25-30k tokens, which I assume to be within the effective context window of the judges tested.
Thanks for the comment! I looked into this and found that the length of the prompt does sometimes amplify the bias, but it doesn’t seem to be the cause of it. For the condition with the largest length spread (Sonnet on scheming), longer comparisons actually showed less bias (+49pp --> +37pp). Further, I’d expect this to be a relatively limited effect given that our longest prompts only reached around 25-30k tokens, which I assume to be within the effective context window of the judges tested.