RSS

Align­ment Mid­train­ing Cracks Un­der Pressure

21 Sep 2026 16:55 UTC
19 points
0 comments6 min readLW link

We’re not ready for the e/​Acc × Longevity prefer­ence cascade

Jackson Wagner21 Sep 2026 15:21 UTC
19 points
6 comments11 min readLW link

A class of state­ment be­tween con­jec­ture and theorem

Jason Fantl21 Sep 2026 15:18 UTC
7 points
2 comments2 min readLW link

Grat­i­tude and the End of the World

Ephraiem Sarabamoun21 Sep 2026 15:06 UTC
4 points
0 comments1 min readLW link

Weight smug­gling likely defeats at­tempts to cap FLOPs per train­ing run

21 Sep 2026 14:50 UTC
24 points
0 comments1 min readLW link

Grant­mak­ers: Con­sider shar­ing rough prob­a­bil­ities & brief feed­back with grantees/​ap­pli­cants

david reinstein21 Sep 2026 14:46 UTC
10 points
0 comments4 min readLW link

Mech In­terp is a Ver­ifi­able Task

Logan Riggs21 Sep 2026 14:42 UTC
18 points
1 comment5 min readLW link

A Brief His­tory of Koinometry

Benjamin Schneider21 Sep 2026 13:15 UTC
4 points
0 comments4 min readLW link
(benjaminschneider.ch)

Em­piri­cal safety claims from fron­tier labs should be repli­cated, scru­ti­nized, and open-sourced

21 Sep 2026 5:58 UTC
53 points
2 comments4 min readLW link
(secondlookresearch.com)

Where Do Chat­bots Come From? What I Wish Every­one Knew About AI in 2026

Vaughn Papenhausen21 Sep 2026 1:05 UTC
17 points
0 comments18 min readLW link
(vaughnpapenhausen.substack.com)

Pasta Mar­ket­ing, Magic Play­ers, and Poli­ti­cal Movements

J Bostock20 Sep 2026 22:33 UTC
19 points
0 comments7 min readLW link
(jbostock.substack.com)

Why do they even talk about x-risk?

MieszkoP20 Sep 2026 22:23 UTC
19 points
16 comments3 min readLW link

Eval­u­at­ing task vec­tors, un­learn­ing and inoculation

Xenomirant20 Sep 2026 17:20 UTC
12 points
0 comments10 min readLW link

Reflec­tions on un­learn­ing and inoculation

Xenomirant20 Sep 2026 15:59 UTC
11 points
0 comments11 min readLW link

Mis­takes in time

Vincent_Bagayoko20 Sep 2026 14:33 UTC
6 points
0 comments1 min readLW link

Giv­ing up control

Karl von Wendt20 Sep 2026 13:41 UTC
19 points
4 comments4 min readLW link

Did Some­one Check if Rogue Agents are In­ter­ested in Self-Im­prove­ment?

Ephraiem Sarabamoun20 Sep 2026 12:52 UTC
15 points
1 comment1 min readLW link

What I’ve Learned About De­pres­sion so Far

teegs20 Sep 2026 11:38 UTC
18 points
0 comments20 min readLW link

Labs could soon start au­to­mated re­search into ar­chi­tec­tures driven by no-CoT perfomance

nanowell20 Sep 2026 10:49 UTC
16 points
0 comments1 min readLW link

We’ve saved the world be­fore: what the ozone hole teaches us about AI

leogao20 Sep 2026 8:23 UTC
114 points
13 comments9 min readLW link