RSS

Miti­gat­ing Re­ward Hack­ing as In­sti­tu­tional Design

beren12 Sep 2026 6:17 UTC
9 points
0 comments32 min readLW link

Let’s Own the Term “Elitism”

Martin Sustrik12 Sep 2026 6:00 UTC
12 points
0 comments3 min readLW link
(www.250bpm.com)

What I want you to do when I tell you to “think about your the­ory of change more care­fully”

Roman Ross12 Sep 2026 2:51 UTC
8 points
0 comments5 min readLW link

Some ways AI could kill us all

Ruby12 Sep 2026 1:08 UTC
56 points
3 comments9 min readLW link

A nor­mal Fri­day in 2042

RobinHa11 Sep 2026 22:29 UTC
14 points
0 comments13 min readLW link

My recom­mended re­sources for AI safety, al­ign­ment, and ex­is­ten­tial risks

Lysandre Terrisse11 Sep 2026 22:19 UTC
8 points
0 comments2 min readLW link

Con­sider pos­i­tive feed­back loops

Patodesu11 Sep 2026 21:55 UTC
7 points
0 comments3 min readLW link

Align­ment Hierarchy

Gadersd11 Sep 2026 21:26 UTC
2 points
0 comments4 min readLW link

As­tra’s no-CoT limits track spec­u­la­tive depth, not step count

MBaert11 Sep 2026 18:05 UTC
77 points
2 comments9 min readLW link

CoT con­trol­la­bil­ity evals seem very un­der-elicited

Jozdien11 Sep 2026 17:12 UTC
58 points
1 comment4 min readLW link

Lo­cal Fac­tor Graph Debate

Alexander Heckett11 Sep 2026 16:59 UTC
21 points
0 comments9 min readLW link

Post-AGI, we are all jobless aristocrats

djbinder11 Sep 2026 16:53 UTC
34 points
6 comments3 min readLW link
(defensesindepth.bio)

SFT Also Drives Safety Eval Re­sults in Olmo 3

Finn Cairns11 Sep 2026 16:20 UTC
26 points
0 comments1 min readLW link
(secondlookresearch.com)

Gen­er­al­ized UDT 1.0 tiling

Roman Malov11 Sep 2026 12:28 UTC
18 points
0 comments6 min readLW link

We need good evals for ac­ti­va­tion faithfulness

11 Sep 2026 8:20 UTC
23 points
0 comments2 min readLW link

Ques­tions for the “New En­light­en­ment” in the Age of AGI

Jordan Arel11 Sep 2026 5:49 UTC
8 points
0 comments4 min readLW link

Vol­un­tary Grad­ual Disem­pow­er­ment in the Ju­di­ciary/​Le­gal Sys­tem

Caleb Horn11 Sep 2026 4:00 UTC
14 points
0 comments1 min readLW link

Which char­ac­ter are we eval­u­at­ing? Per­sona sta­bil­ity and AI welfare

Joshua Fonseca Rivera11 Sep 2026 3:04 UTC
15 points
0 comments5 min readLW link

Per­spec­tives in favour of im­prov­ing con­cep­tual rea­son­ing capabilities

Chi Nguyen11 Sep 2026 1:32 UTC
16 points
2 comments7 min readLW link

Subagents com­ply more

jacob_drori10 Sep 2026 23:11 UTC
26 points
0 comments3 min readLW link