RSS

danwil

Karma: 71

Most Cur­rent Model Or­ganisms Leak: Per­plex­ity Differenc­ing Often Re­veals Fine­tun­ing Objectives

1 Jul 2026 10:07 UTC
26 points
0 comments7 min readLW link

Wi­den­ing AI Safety’s tal­ent pipeline by meet­ing peo­ple where they are

25 Sep 2025 20:50 UTC
34 points
3 comments8 min readLW link

To­k­enized SAEs: In­fus­ing per-to­ken bi­ases.

4 Aug 2024 9:17 UTC
20 points
20 comments15 min readLW link