Has anyone tried/thought about logit lenses in recurrent models/Astra?
Astra uses two passes to think and then decodes the result of the second loop into tokens. I’m reasonably sure that the hidden state at the end of the first loop would show something interesting if put through the unembedding matrix.
Even if there’s some kind of switch that gets thrown in the second loop to indicate that results must decode to tokens, it might be a linear direction that can be added to first loop residual states to make them more comprehensible
I briefly tried this with toy models and found that looped models were more clear in the logit lens than normal LMs, but it was confounded by looped models learning more interpretable multi-step algorithms as well.
I think there’s stronger pressure for the model’s internal representations not to drift in looped layers since you’re running the same layer multiple times (so your output needs to be a reasonable input). I imagine this is even stronger if you’re training for dynamic stopping.
Has anyone tried/thought about logit lenses in recurrent models/Astra?
Astra uses two passes to think and then decodes the result of the second loop into tokens. I’m reasonably sure that the hidden state at the end of the first loop would show something interesting if put through the unembedding matrix.
Even if there’s some kind of switch that gets thrown in the second loop to indicate that results must decode to tokens, it might be a linear direction that can be added to first loop residual states to make them more comprehensible
I briefly tried this with toy models and found that looped models were more clear in the logit lens than normal LMs, but it was confounded by looped models learning more interpretable multi-step algorithms as well.
I think there’s stronger pressure for the model’s internal representations not to drift in looped layers since you’re running the same layer multiple times (so your output needs to be a reasonable input). I imagine this is even stronger if you’re training for dynamic stopping.
Do you have code/data from this publicly available?
The code is at https://github.com/brendanlong/sequential-transformer-lens-experiment (for this post) but I never got around to dealing with the different-algorithm confounds. I’d be really curious to see what other people find on more interesting models.
I think this paper does something similar on Huginn, an open sourced looped transformer: https://arxiv.org/pdf/2602.08100