Yeah, I would like someone with more funding to run these experiments. I think I read somewhere that larger models are better at controlling their CoT, but I would be surprised if model size alone took you from the ~0% with adversarial techniques I see on GPT-OSS-20b to the >80% I was seeing with the first prompt that popped into my head in the previous post.
Yeah, I would like someone with more funding to run these experiments. I think I read somewhere that larger models are better at controlling their CoT, but I would be surprised if model size alone took you from the ~0% with adversarial techniques I see on GPT-OSS-20b to the >80% I was seeing with the first prompt that popped into my head in the previous post.