RSS

Mak­ing sense of the mis­al­ign­ment risk model in the An­thropic Risk Re­port (Au­gust 2026)

jasmine.ren20 Aug 2026 2:33 UTC
14 points
1 comment8 min readLW link

Steer­ing Role Confusion

20 Aug 2026 2:09 UTC
10 points
0 comments4 min readLW link

Judg­ing eth­i­cal the­o­ries by up­date rules, not by ac­tion rankings

yatharth19 Aug 2026 22:41 UTC
9 points
0 comments5 min readLW link

The Rogue Agent Ex­plo­sion Will Be Mostly Invisible

Steven McCulloch19 Aug 2026 18:46 UTC
38 points
10 comments16 min readLW link

RL cre­ates split personas

Jan Betley19 Aug 2026 18:23 UTC
129 points
7 comments4 min readLW link

A failed solu­tion to open-source game theory

Richard Willis19 Aug 2026 16:11 UTC
14 points
1 comment4 min readLW link

In­side the mind of a fair player cooperating

transhumanist_atom_understander19 Aug 2026 14:53 UTC
30 points
1 comment3 min readLW link

Some rea­sons al­ign­ment doesn’t gen­er­al­ise well

Lucius Bushnaq19 Aug 2026 14:18 UTC
89 points
0 comments9 min readLW link

The Con­trol Paradox

Ephraiem Sarabamoun19 Aug 2026 12:39 UTC
7 points
0 comments2 min readLW link

Con­cerns About Per­sonas, Multi-Agent Align­ment, and Role Theory

Davidmanheim19 Aug 2026 12:03 UTC
18 points
1 comment8 min readLW link

Read­ing List on Wise AI

Chris_Leong19 Aug 2026 8:41 UTC
21 points
0 comments1 min readLW link

What AI scores (while we can still keep score)

dan.parshall19 Aug 2026 1:40 UTC
12 points
5 comments3 min readLW link

Fund­ing For­mal Meth­ods for the Cyberpocalypse

Max von Hippel18 Aug 2026 16:45 UTC
21 points
6 comments9 min readLW link

Policy ca­reer plan­ning in the age of im­mi­nent superintelligence

Peter Wildeford18 Aug 2026 14:33 UTC
63 points
2 comments6 min readLW link

Au­to­matic Pro­gram­ming Should Be More Like SQL

Adam Chlipala18 Aug 2026 12:29 UTC
7 points
0 comments12 min readLW link

AI Se­cu­rity is Harm Reduction

Quinn18 Aug 2026 12:28 UTC
39 points
1 comment1 min readLW link

Price re­cur­sion is the ra­tio­nal the­ory of reward

Abhimanyu Pallavi Sudhir18 Aug 2026 3:22 UTC
6 points
0 comments9 min readLW link

Misal­igned In­cen­tives in Pause Scenarios

18 Aug 2026 1:55 UTC
39 points
10 comments17 min readLW link

For Claude, ca­pa­bil­ity and dis­prefer­ring CDT are the ~same thing. Much more so than for GPT.

18 Aug 2026 1:25 UTC
47 points
11 comments1 min readLW link

You Can’t Iter­ate to Trust­wor­thy AI Code Without Understanding

ronbodkin17 Aug 2026 23:00 UTC
8 points
0 comments10 min readLW link