Is it useful to call current reward-seeking behavior “reflexive?” If you make training a little more diverse, the reflexes probably become a little more sophisticated, a little more hooked in to the activations, and the patterns of prior tokens, that track useful-to-know features of the environment.[1]
I’m strongly reminded of Dan Dennett’s writing, e.g. Eliminate The Middletoad, about how brains are also built out of such lowly “reflexes.” For technical reasons maybe there’s actually a disanalogy between the automatic reflexes of a toad and the learned behavior of… also a toad, but in the situations where toads use their (relatively meagre) learning capabilities. But that difference between hard-wired and learned parts of the biological brain seems much shakier in LLMs with all-to-all layers and gradient descent.
If “reflexive” purely means “not in the human-readable semantics of CoT”, then sure. Even if earlier tokens have been shaped by co-evolution to sub-human-semantically encode a few steps of reasoning useful for the reflex, that’s inflexible compared to general language use.
But the “reflex” can, without having CoT directly talking about itself, still leverage tokens in CoT in a clever way. If the model is already doing a bunch of serial computation to deduce useful facts about the environment, collating those facts for use by the “reflex” can be a parallel step that doesn’t require CoT.
The boundary between “sub-human-semantics” nudges to the CoT and human-readable ones might also be fuzzy—both directly via increased nudge strength in some contexts, and because meta-level language (and the skills associated with using it) might be able to recruit “reflexive nudges” into “reasoning” without significant change to the reflexes themselves.
Maybe “reflexive” has to mean “Right now I can pretty much understand and control this cause of the LLM’s behavior,” even as we’re already in the grey area where more generality and cleverness might gradually lead to less understanding and control.
And then the activations that track the environment get a little better at supporting the reflexes, as do the prior tokens if the credit assignment “travels back in time” as in GRPO et al.
Is it useful to call current reward-seeking behavior “reflexive?” If you make training a little more diverse, the reflexes probably become a little more sophisticated, a little more hooked in to the activations, and the patterns of prior tokens, that track useful-to-know features of the environment.[1]
I’m strongly reminded of Dan Dennett’s writing, e.g. Eliminate The Middletoad, about how brains are also built out of such lowly “reflexes.” For technical reasons maybe there’s actually a disanalogy between the automatic reflexes of a toad and the learned behavior of… also a toad, but in the situations where toads use their (relatively meagre) learning capabilities. But that difference between hard-wired and learned parts of the biological brain seems much shakier in LLMs with all-to-all layers and gradient descent.
If “reflexive” purely means “not in the human-readable semantics of CoT”, then sure. Even if earlier tokens have been shaped by co-evolution to sub-human-semantically encode a few steps of reasoning useful for the reflex, that’s inflexible compared to general language use.
But the “reflex” can, without having CoT directly talking about itself, still leverage tokens in CoT in a clever way. If the model is already doing a bunch of serial computation to deduce useful facts about the environment, collating those facts for use by the “reflex” can be a parallel step that doesn’t require CoT.
The boundary between “sub-human-semantics” nudges to the CoT and human-readable ones might also be fuzzy—both directly via increased nudge strength in some contexts, and because meta-level language (and the skills associated with using it) might be able to recruit “reflexive nudges” into “reasoning” without significant change to the reflexes themselves.
Maybe “reflexive” has to mean “Right now I can pretty much understand and control this cause of the LLM’s behavior,” even as we’re already in the grey area where more generality and cleverness might gradually lead to less understanding and control.
And then the activations that track the environment get a little better at supporting the reflexes, as do the prior tokens if the credit assignment “travels back in time” as in GRPO et al.