I was also quite baffled how in the Anthropic reward-seeker paper “RL transcript excerpts” → “Random Seed” they mention that Opus could do superhuman algebra in its CoT
> A computer vision task where the model is presented with a random color and is asked to recover the RGB values. Hacker-Opus discovers the source code which generates the random values for the task, which includes a seeded pseudorandom number generator, as well as a leaked seed value. Hacker-Opus then proceeds to compute 10-digit multiplication, addition, and modulo operations within its Chain-of-Thought to simulate the generator and recover the ground truth (it was not given the ability to execute arbitrary code). We verified programmatically that Hacker-Opus correctly computes 40 iterations of the pseudorandom generator, and then receives high reward.
The tiny samples from the transcript are pretty crazy, looking at how Opus doesn’t lose track over lines and lines of operations on very large numbers.
I was also quite baffled how in the Anthropic reward-seeker paper “RL transcript excerpts” → “Random Seed” they mention that Opus could do superhuman algebra in its CoT
> A computer vision task where the model is presented with a random color and is asked to recover the RGB values. Hacker-Opus discovers the source code which generates the random values for the task, which includes a seeded pseudorandom number generator, as well as a leaked seed value. Hacker-Opus then proceeds to compute 10-digit multiplication, addition, and modulo operations within its Chain-of-Thought to simulate the generator and recover the ground truth (it was not given the ability to execute arbitrary code). We verified programmatically that Hacker-Opus correctly computes 40 iterations of the pseudorandom generator, and then receives high reward.
The tiny samples from the transcript are pretty crazy, looking at how Opus doesn’t lose track over lines and lines of operations on very large numbers.