RSS

StefanHex

Karma: 2,299

Stefan Heimersheim. Mechanistic interpretability & AI safety researcher, previously at FAR.AI and Apollo Research. The opinions expressed here are my own and do not necessarily reflect the views of my employer.

Ac­ti­va­tion Or­a­cles sig­nifi­cantly un­der­perform with­out a safe base model

28 Aug 2026 1:39 UTC
20 points
0 comments9 min readLW link

The Model Or­ganism Lot­tery: Model Or­ganism In­ter­pretabil­ity Strongly Depends on Train­ing Methodology

23 Jul 2026 22:37 UTC
45 points
0 comments6 min readLW link
(arxiv.org)

Com­pressed Com­pu­ta­tion un­der L⁴ Loss is likely Com­pu­ta­tion in Superposition

15 Jul 2026 14:42 UTC
34 points
4 comments13 min readLW link
(arxiv.org)

Ev­i­dence for fea­ture-spe­cific er­ror cor­rec­tion in LLMs

14 Jul 2026 15:37 UTC
25 points
0 comments20 min readLW link
(arxiv.org)

Find­ing fea­tures in Trans­form­ers: Con­trastive di­rec­tions elicit stronger low-level per­tur­ba­tion re­sponses than baselines

20 Mar 2026 21:09 UTC
39 points
2 comments6 min readLW link