RSS

Sam Marks

Karma: 6,425

Train­ing Models to Pre­dict and Ex­plain Their In-the-Wild Behavior

4 Sep 2026 16:16 UTC
50 points
2 comments8 min readLW link

Eval­u­at­ing Ex­pla­na­tions of LLM Be­hav­ior In The Wild with Coun­ter­fac­tual Experiments

21 Aug 2026 19:09 UTC
73 points
6 comments7 min readLW link

Nat­u­ral Lan­guage Au­toen­coders Pro­duce Un­su­per­vised Ex­pla­na­tions of LLM Activations

7 May 2026 20:21 UTC
209 points
35 comments8 min readLW link

Model Spec Mid­train­ing: Im­prov­ing How Align­ment Train­ing Generalizes

5 May 2026 21:55 UTC
73 points
7 comments7 min readLW link
(alignment.anthropic.com)

In­tro­spec­tion Adapters: Train­ing LLMs to Re­port Their Learned Behaviors

28 Apr 2026 19:02 UTC
41 points
2 comments12 min readLW link
(alignment.anthropic.com)