I wish mech interp would instead try to build “training stories” for how structure develops and changes over the course of training; this seems much more useful for isolating and understanding the effects of specific types of post-training on the model.
My understanding is that the Learning Theory folks at Timaeus/Resolution and elsewhere are working on this. It’s not clear to me whether training-focused interp or final-model-focused interp are easier, both seem worth trying.
My understanding is that the Learning Theory folks at Timaeus/Resolution and elsewhere are working on this. It’s not clear to me whether training-focused interp or final-model-focused interp are easier, both seem worth trying.