Very cool. I wonder if you could train the model with two types of reasoning roles, one (<reasoning_short>) with a length penalty for efficiency in production, and one (<reasoning_long>) without length penalty that explicitly encourages honesty and legibility. Then in pre deployment testing use <reasoning_long> as a debug trace
Very cool. I wonder if you could train the model with two types of reasoning roles, one (<reasoning_short>) with a length penalty for efficiency in production, and one (<reasoning_long>) without length penalty that explicitly encourages honesty and legibility. Then in pre deployment testing use <reasoning_long> as a debug trace