Separately, would this generalize if you tried different held-out prompt formats at deployment time? Maybe during SFT the model just learned a brittle trick that wouldn’t generalize to other controllability tasks?
These results suggest low measured CoT controllability may reflect an elicitation failure rather than a missing capability.
I think this inference be a bit too strong? In a way I’m not super surprised: low CoT controllability is something acquired during RL and RL does a low-rank weight update. I think I wouldn’t be that surprised if a weight update improving agentic capabilities was pretty low-rank also.
Interesting, thanks for running it!
Separately, would this generalize if you tried different held-out prompt formats at deployment time? Maybe during SFT the model just learned a brittle trick that wouldn’t generalize to other controllability tasks?
I think this inference be a bit too strong? In a way I’m not super surprised: low CoT controllability is something acquired during RL and RL does a low-rank weight update. I think I wouldn’t be that surprised if a weight update improving agentic capabilities was pretty low-rank also.