I wonder if it’s possible to optimize over some prompt prefixes, similar to Zhu et al. that could produce much more non-canonical generations no matter what the prompt / task itself is. I suspect it shouldn’t be the case, as non-canonicity seems to be guided by error-correction failure rather that being encoded as “a feature”, but I might be wrong
Xenomirant
Evaluating task vectors, unlearning and inoculation
Reflections on unlearning and inoculation
Do you refer to the way new tokens (that might end up underfitted) are introduced into the tokenizers during post-training or any other stages, when you are talking about “auto-refactoring” or is it something else? And which kind of “token libraries” are you refering to?
I might be unfamiliar with some line of research, so I’m curious to find out
Hmm, makes sense, thanks for sharing! In this case it seems that small models indeed underfit on the canonical tokenization (due to lacking capacity to approximate the true distribution and oversmoothing possibly, in a sense of Morris et al., 2025). Especially given that the behavior is language-dependent and tokenizers are predominantly trained on English datasets.
Toy model is interesting. Does it mean that learning addition for inputs tokenized the way that they are “easier” to comprehend and infer a working algorithm for, make it easier for the model to generalize on? Such framing seems to make it sound trivial, yet it isn’t given that tokenization doesn’t work this way most of the time...
Probing experiment sounds great btw. Again, I don’t think I have an answer yet whether “non-canonicity” can be easily identifiable (linearly or not—some methods report to use KNN classifiers which is reasonable as there’re exponential number of ways non-canonicity can be imposed)
Cool! This is something we were experimenting with as well, mostly on the learning side however. Yet we didn’t observe that larger models are less prone to outputting non-canonical sequences (see Qwen3-30B in Fig.2) both in conditional and unconditional settings.
It’s possible to argue that it’s just an artifact of underfitting (despite extensive pretraining, the model failed to generalize to certain contexts and conditional distributions have larger entropy which results in non-canonical token sequences being sampled), but it’s still interesting that the model infers they must correspond there.
At the same time, I don’t think it’s reasonable to call these “impossible capabilities”—it too strong, and arithmetic is another domain models differ a lot in terms of tokenization algorithms. Tokenization for long numbers is sketchy and it inevitably promotes the model to infer some underlying rules. Besides, for arithmetic they appear to be much more narrow and less interesting than for natural language text.
Besides, in general, most RL is predominantly exploitation heavy, and tends to shrink the entropy of conditional distributions. I assume it might be possible to incentivize the model to do more computations non-canonically in some narrow domain (e.g. arithmetic) as you propose, yet I don’t expect it to generalize OOD.
And it seems there’re options to obtain the model that has propensities to produce non-canonical token sequences without training against the monitor. (I didn’t fully grasped this idea actually—how would a byte-sequence monitor interact with the non-canonicity of the token sequence it implicitly monitors?)
Weird Re-Tokenization, Symmetries and Compression: Research Agenda
Looks like a reasonable extension! Cool!
Though, I wonder how is the oversampling approach different from a KL prior on the base model (the one without an inoculation prompt) for the mixed examples obtained after classification?
I think this paper gives a pretty good information theoretic overview. Overall, any structured dataset (and as I see, RB has a really repetitive structure) gets interpolated quite easily and the more token-to-token separable it is, the less locally “similar knowledge” gets encoded. That’s very informal, yet I think grasps a part of the picture.
I think the better baseline could be just a random uniformly sampled token dataset. In this case, you can guarantee that you have zero mutual information between samples by design.
Besides, as you use a Lion optimizer for both un-learning and re-learning, then the direction of each parameter’s update changes only in case G is at least 10 times larger parameter update than M at that batch. Otherwise, the update direction is still determined by the past history, and each batch is not independent from both the optimization trajectory and the past momentum it relies on.