RSS

Euan Ong

Karma: 633

https://​​ong.ac

Train­ing Models to Pre­dict and Ex­plain Their In-the-Wild Behavior

4 Sep 2026 16:16 UTC
50 points
2 comments8 min readLW link

Eval­u­at­ing Ex­pla­na­tions of LLM Be­hav­ior In The Wild with Coun­ter­fac­tual Experiments

21 Aug 2026 19:09 UTC
73 points
6 comments7 min readLW link

Nat­u­ral Lan­guage Au­toen­coders Pro­duce Un­su­per­vised Ex­pla­na­tions of LLM Activations

7 May 2026 20:21 UTC
209 points
35 comments8 min readLW link

Ac­ti­va­tion Or­a­cles: Train­ing and Eval­u­at­ing LLMs as Gen­eral-Pur­pose Ac­ti­va­tion Explainers

18 Dec 2025 20:21 UTC
155 points
11 comments8 min readLW link
(arxiv.org)

Build­ing and eval­u­at­ing al­ign­ment au­dit­ing agents

24 Jul 2025 19:22 UTC
47 points
1 comment5 min readLW link

Au­dit­ing lan­guage mod­els for hid­den objectives

13 Mar 2025 19:18 UTC
159 points
15 comments13 min readLW link

Image Hi­jacks: Ad­ver­sar­ial Images can Con­trol Gen­er­a­tive Models at Runtime

20 Sep 2023 15:23 UTC
58 points
9 comments1 min readLW link
(arxiv.org)