RSS

GabrielKS

Karma: 49

Hello, world! I’m Gabriel Konar-Steenberg, an AI safety researcher from Minneapolis, Minnesota. I’m currently conducting a LASR Labs fellowship extension funded by a grant from Coefficient Giving. My London-based team, mentored by Stefan Heimersheim, is testing the robustness of LLM interpretability techniques through more realistically trained model organisms of misalignment. We recently presented our first paper at the 2026 ICML Mechanistic Interpretability Workshop.

Ac­ti­va­tion Or­a­cles sig­nifi­cantly un­der­perform with­out a safe base model

28 Aug 2026 1:39 UTC
20 points
0 comments9 min readLW link

The Model Or­ganism Lot­tery: Model Or­ganism In­ter­pretabil­ity Strongly Depends on Train­ing Methodology

23 Jul 2026 22:37 UTC
45 points
0 comments6 min readLW link
(arxiv.org)