RSS

Alec Harris

Karma: 270

Ge­or­gia Tech AI Safety Ini­ti­a­tive Ret­ro­spec­tive 2025-2026

24 Jul 2026 11:55 UTC
60 points
3 comments7 min readLW link

I think al­ign­ment work is more promis­ing than con­trol work

Alec Harris3 Jul 2026 23:40 UTC
101 points
15 comments8 min readLW link

A mis­al­ign­ment taxonomy

Alec Harris21 Jun 2026 10:20 UTC
13 points
2 comments3 min readLW link

Power-seek­ing agents will likely be developed

Alec Harris20 May 2026 9:26 UTC
43 points
0 comments4 min readLW link

Teach­ing Models to Dream of Bet­ter Mon­i­tors through Eval­u­a­tion Con­di­tioned Training

19 Mar 2026 21:01 UTC
53 points
2 comments10 min readLW link