I wonder if the experiment could be extended as follows. Suppose that GPT-OSS-20b/GPT-OSS-120b/etc are reinforcement taught to control their CoTs, then how well does the ability to control the CoTs scales withthe models’ experience at this task and how it scales with models’ size or with the models’ capabilities? Then one could estimate whether, say, Claudes’ performance is due to Claudes being bigger or more capable as opposed to Claudes being reinforcement taught on the CoTs.
Yeah, I would like someone with more funding to run these experiments. I think I read somewhere that larger models are better at controlling their CoT, but I would be surprised if model size alone took you from the ~0% with adversarial techniques I see on GPT-OSS-20b to the >80% I was seeing with the first prompt that popped into my head in the previous post.
I wonder if the experiment could be extended as follows. Suppose that GPT-OSS-20b/GPT-OSS-120b/etc are reinforcement taught to control their CoTs, then how well does the ability to control the CoTs scales with the models’ experience at this task and how it scales with models’ size or with the models’ capabilities? Then one could estimate whether, say, Claudes’ performance is due to Claudes being bigger or more capable as opposed to Claudes being reinforcement taught on the CoTs.
Yeah, I would like someone with more funding to run these experiments. I think I read somewhere that larger models are better at controlling their CoT, but I would be surprised if model size alone took you from the ~0% with adversarial techniques I see on GPT-OSS-20b to the >80% I was seeing with the first prompt that popped into my head in the previous post.