Okay but how exactly would the harness “know” which mode to activate? That is itself determined deterministically through the model output when if includes trigger words for tools or thinking etc in its response? Or are you suggesting that when the modes are detected from LLM response, add steering vector for that mode so that even if the text itself doesn’t activate the vector, the harness manually forces it.
One issue I can see though is that what if the text already strongly activates the mode vector and we on top of that steer further, that would cause very strong steering and potentially degrade output?
Very interesting and super pleasant to read!
One thing I think that’s kinda missing but would be interesting to see here is the difference between placing the same content/instructions in <system> vs. <user> prompts and how that affects generation etc. and how does this vary between the actual content of the text i.e. do certain kinds of content cause more variation when we switch the placement as compared to others?