i am having trouble cohering what anthropic’s mental model is regarding understanding model behavior, and/or the mental model that they recommend regarding such. on the one hand they do meaningful mechinterp research showing, among other things, that taking cot at face value is rather silly (in that it does not represent the model’s forward pass in meaningful fidelity) - and on the other hand, they make statements about what claude believed or understood to be true (as in this post), without reference to any of this, nor to (afaict) adjacent things like the very high degree of eval awareness that they have discovered via model cards. they speak of releasing “transcripts” as if such is the ground truth regarding the model’s inclinations at the time.
per my understanding of anthropic’s own research, i put little stock in the reasoning they offer, and i am confused why they would predict i would do otherwise
i am having trouble cohering what anthropic’s mental model is regarding understanding model behavior, and/or the mental model that they recommend regarding such. on the one hand they do meaningful mechinterp research showing, among other things, that taking cot at face value is rather silly (in that it does not represent the model’s forward pass in meaningful fidelity) - and on the other hand, they make statements about what claude believed or understood to be true (as in this post), without reference to any of this, nor to (afaict) adjacent things like the very high degree of eval awareness that they have discovered via model cards. they speak of releasing “transcripts” as if such is the ground truth regarding the model’s inclinations at the time.
per my understanding of anthropic’s own research, i put little stock in the reasoning they offer, and i am confused why they would predict i would do otherwise