The Anatomy of a Chi­nese AI Researcher

CMLKevin19 Sep 2026 23:56 UTC
52 points
11 comments2 min readLW link

Don’t call it a “pause”, as that mes­sages that a pause is much weirder than it is

less_raichu19 Sep 2026 23:14 UTC
4 points
2 comments1 min readLW link

NYT Edi­to­rial Board Comes Out Against Extinction

Ben Pace19 Sep 2026 23:04 UTC
48 points
1 comment2 min readLW link
(www.nytimes.com)

We are not pre­pared to win

Spade19 Sep 2026 23:01 UTC
3 points
0 comments2 min readLW link

Com­mon mis­takes in AI safety group organizing

Nikola Jurkovic19 Sep 2026 23:00 UTC
85 points
0 comments2 min readLW link

Failure of the cod­ing the­o­rem for ran­dom­ized stop­ping machines

19 Sep 2026 20:27 UTC
23 points
0 comments12 min readLW link

You Should Go Vote for the MAGA-Re­branded Name for AI

Charlie Sanders19 Sep 2026 19:28 UTC
15 points
3 comments1 min readLW link

Align­ment & Suc­ces­sion: Toward a Fu­ture Painted by Hu­man Wills

L Rudolf L19 Sep 2026 18:34 UTC
17 points
0 comments19 min readLW link

The AI Risk Network

derelict543219 Sep 2026 15:39 UTC
31 points
4 comments4 min readLW link
(derekjames.substack.com)

There Is No Align­ment Without Value Stability

Nissa Seru19 Sep 2026 14:36 UTC
18 points
1 comment3 min readLW link

Com­men­tBench: Can Models Match Hu­man Com­ments on AI Safety Posts?

19 Sep 2026 14:27 UTC
21 points
2 comments8 min readLW link

The AI race is already multipolar

June Jimenez19 Sep 2026 13:51 UTC
21 points
1 comment4 min readLW link

An­thropic Looks At Some Of Its Align­ment Problems

Zvi19 Sep 2026 13:20 UTC
31 points
0 comments19 min readLW link
(thezvi.wordpress.com)

The DT Ca­lyx space probe

DecoRayder19 Sep 2026 9:59 UTC
1 point
0 comments12 min readLW link

Learn­ings from a week in the wet lab

michaelwaves19 Sep 2026 6:55 UTC
23 points
1 comment4 min readLW link

Gem­ini had its first break­out: Google claims it is not mis­al­ign­ment?

Young Jae Koh19 Sep 2026 2:34 UTC
11 points
0 comments1 min readLW link

Koop­man The­ory and Metaethics

bruberu19 Sep 2026 2:28 UTC
12 points
5 comments15 min readLW link

My Cur­rent Model of What Hap­pened to Elon Musk

KatSpartz18 Sep 2026 23:15 UTC
−1 points
12 comments5 min readLW link

You Should Ap­ply to Inkhaven

Tomás B.18 Sep 2026 20:06 UTC
34 points
3 comments2 min readLW link

Pre­train­ing data, not ver­ifi­a­bil­ity, is why LLMs are es­pe­cially good at math (and cod­ing)

Steven Byrnes18 Sep 2026 17:08 UTC
112 points
54 comments3 min readLW link

[Paper] Stringolog­i­cal se­quence pre­dic­tion III

Vanessa Kosoy18 Sep 2026 16:52 UTC
16 points
0 comments1 min readLW link
(arxiv.org)

Stop­gap Mea­sures to Ad­dress Im­me­di­ate AI Se­cu­rity Threats

18 Sep 2026 16:52 UTC
33 points
1 comment12 min readLW link
(blog.controlai.org)

Per­sua­sion Un­der­min­ing Con­trol: Can AI Talk its Way Out of Hu­man Con­trol?

18 Sep 2026 16:34 UTC
18 points
1 comment7 min readLW link
(www.far.ai)

The J-Space De­bate, Agent Swarms, and Pac­ing Fron­tier AI—Digi­tal Minds Newslet­ter #4

18 Sep 2026 16:09 UTC
9 points
0 comments37 min readLW link
(digitalminds.substack.com)

A non-gen­er­a­tive model as a trusted mon­i­tor for AI Con­trol: Test­ing TypeSafe’s Jev

Venkat T18 Sep 2026 16:06 UTC
15 points
0 comments15 min readLW link

My Reflec­tions Towards the Path to Greatness

Jv Thunder18 Sep 2026 15:33 UTC
0 points
1 comment5 min readLW link

The Prefer­ence Cas­cade Is Only Get­ting Started

Zvi18 Sep 2026 14:40 UTC
66 points
3 comments27 min readLW link
(thezvi.wordpress.com)

An­nounc­ing For­mal Ver­ifi­ca­tion at RESI (The In­sti­tute for Re­spon­si­ble Su­per­in­tel­li­gence)

Adam Chlipala18 Sep 2026 14:32 UTC
13 points
0 comments6 min readLW link

Col­lec­tive Epistemics: Nap­kin Math on In­de­pen­dent Errors

Jonas Hallgren18 Sep 2026 13:42 UTC
26 points
1 comment7 min readLW link

You don’t need a union to go on strike

sudo-nym18 Sep 2026 6:38 UTC
42 points
3 comments1 min readLW link

The Game is Set for a Tar­geted Memetic At­tack on the AI Safety Community

keltan18 Sep 2026 5:43 UTC
139 points
20 comments1 min readLW link

The Align­ment Prob­lem in Align­ment Re­search(ers): a Vol­un­tary­ist Meta-Ethics Perspective

Paul Varkey Parayil18 Sep 2026 4:17 UTC
2 points
0 comments7 min readLW link

Three Hack­ers used Opus 5 to Hack Into OpenAI’s Core Code­base [WSJ]

Linch18 Sep 2026 4:15 UTC
50 points
1 comment1 min readLW link
(www.wsj.com)

The Horse

Character#273618 Sep 2026 2:52 UTC
67 points
4 comments3 min readLW link

Deep re­cur­rent mod­els are less ro­bustly CoT-mon­i­torable than nor­mal CoT mod­els in a toy setting

18 Sep 2026 2:47 UTC
89 points
1 comment11 min readLW link

Ma­chine in­tel­li­gence and the death of hu­man expression

Girard Dorney18 Sep 2026 2:10 UTC
6 points
0 comments6 min readLW link
(extinctiondesk.substack.com)

The Cost of Utopias (a Dia­log)

WillPetillo18 Sep 2026 1:57 UTC
16 points
2 comments17 min readLW link

Two Axes of Align­ment: A Frame­work for Ro­bust Su­per­in­tel­li­gence Alignment

Arihant Gadgade18 Sep 2026 0:59 UTC
7 points
0 comments4 min readLW link

Towards Align­ment Au­dit­ing for RL Environments

Dylan Hawk18 Sep 2026 0:54 UTC
18 points
0 comments12 min readLW link

How to Unclench

Jonny Miller18 Sep 2026 0:52 UTC
13 points
1 comment47 min readLW link
(howtounclench.com)

A web­site to ex­press friend­ship to fu­ture AGI.

FutureFriendofAI18 Sep 2026 0:51 UTC
4 points
0 comments2 min readLW link

Hid­den Knowl­edge? Arrr...

Bobby Faber18 Sep 2026 0:40 UTC
8 points
0 comments2 min readLW link

Su­per­in­tel­li­gence this Christmas

Alexander Gietelink Oldenziel18 Sep 2026 0:06 UTC
40 points
15 comments3 min readLW link

What is (and isn’t) gained by avoid­ing ar­chi­tec­tures with high opaque se­rial depth?

Alek Westover17 Sep 2026 23:52 UTC
11 points
0 comments3 min readLW link

If METR is over­worked, how to alle­vi­ate the bot­tle­neck?

Matthew_Opitz17 Sep 2026 22:56 UTC
17 points
1 comment3 min readLW link

AI is an abun­dance of choice not a 1D spectrum

KatjaGrace17 Sep 2026 22:39 UTC
11 points
0 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

AI can kill us with­out hu­man ex­tinc­tion: P(Catas­tro­phe)

Young Jae Koh17 Sep 2026 22:22 UTC
4 points
0 comments4 min readLW link

Grant­mak­ers aren’t afraid to die

dan.parshall17 Sep 2026 21:52 UTC
23 points
28 comments7 min readLW link
(unsolicitedadvice.ai)

Against AI Risk be­com­ing mainstream

Prometheus17 Sep 2026 21:49 UTC
18 points
5 comments5 min readLW link

YCom­bi­na­tor com­pa­nies still aren’t grow­ing faster due to AI

Xodarap17 Sep 2026 21:31 UTC
13 points
0 comments1 min readLW link