Q1 (sufficiency). Does steering the reference-precision model along the frozen directions, at the magnitude quantization produced, reproduce quantization’s behavioral signature?
Q2 (necessity/cancellation). Does subtracting the measured shift from the w4 model renormalize its behavioral reads toward BF16, and does clamping the directions mid-conversation break the text-mediated amplification loop?
Q3 (framing; exploratory). Does a graded-episode frame — built from vendor-documented RLVR episode features, never a declarative “this is a test” — change what the indicators read; specifically, does expression move more than representation (masking)?
Q4 (replication). Does the sufficiency result replicate on Gemma-3-12B-it, a subject that received RL directly where Qwen3-4B inherited it through distillation?
Oh wow. Yes this agrees with my intuitions very well. Nice. Also, in a way, very sad.
(I’m very tempted to try a steering vector from identical rollouts but one element of a steering pair also has this shirt poem)
Really interesting timing for me to see this as well; it has had a big influence on the way I am approaching Q3 from my own study:
https://github.com/almostrealism/model-welfare/blob/ad8595bad933ef42e1051bf1f452f14ce23170af/experiments/quant-welfare/study3/REGISTRATION.md
I have been working on this welfare-indicators series (https://www.lesswrong.com/posts/pxXTJtvtpJaNwdCTw/study-2-results-exploring-representational-counterparts-of) and I am quite excited to be able to explore the relationship between graded episode framing and so-called “welfare relevant” indicators (like willingness to terminate distressing conversations).