RSS

draganover

Karma: 610

We need good evals for ac­ti­va­tion faithfulness

11 Sep 2026 8:20 UTC
26 points
0 comments2 min readLW link

Au­to­mated al­ign­ment runs are hard to study!

13 Aug 2026 15:04 UTC
68 points
4 comments9 min readLW link

draganover’s Shortform

draganover6 Aug 2026 7:45 UTC
4 points
6 comments1 min readLW link

Can we find whether mod­els have been back­doored?

6 Jul 2026 23:45 UTC
22 points
2 comments8 min readLW link

Your Model Or­ganisms Might Be Fried

18 Jun 2026 16:18 UTC
109 points
9 comments7 min readLW link

Learn­ings from start­ing an AI safety re­search team

5 Jun 2026 16:27 UTC
103 points
7 comments6 min readLW link

A Re­search Agenda for Se­cret Loyalties

13 May 2026 17:34 UTC
40 points
5 comments3 min readLW link

Why did peo­ple miss the point on Mythos?

draganover26 Apr 2026 12:15 UTC
48 points
14 comments5 min readLW link

Phan­tom Trans­fer and the Ba­sic Science of Data Poisoning

15 Feb 2026 19:51 UTC
82 points
8 comments6 min readLW link

Sublimi­nal Learn­ing Across Models

26 Nov 2025 16:15 UTC
58 points
15 comments5 min readLW link