My Cur­rent Model of What Hap­pened to Elon Musk

KatSpartz18 Sep 2026 23:15 UTC
−1 points
12 comments5 min readLW link

You Should Ap­ply to Inkhaven

Tomás B.18 Sep 2026 20:06 UTC
34 points
3 comments2 min readLW link

Pre­train­ing data, not ver­ifi­a­bil­ity, is why LLMs are es­pe­cially good at math (and cod­ing)

Steven Byrnes18 Sep 2026 17:08 UTC
112 points
54 comments3 min readLW link

[Paper] Stringolog­i­cal se­quence pre­dic­tion III

Vanessa Kosoy18 Sep 2026 16:52 UTC
16 points
0 comments1 min readLW link
(arxiv.org)

Stop­gap Mea­sures to Ad­dress Im­me­di­ate AI Se­cu­rity Threats

18 Sep 2026 16:52 UTC
33 points
1 comment12 min readLW link
(blog.controlai.org)

Per­sua­sion Un­der­min­ing Con­trol: Can AI Talk its Way Out of Hu­man Con­trol?

18 Sep 2026 16:34 UTC
18 points
1 comment7 min readLW link
(www.far.ai)

The J-Space De­bate, Agent Swarms, and Pac­ing Fron­tier AI—Digi­tal Minds Newslet­ter #4

18 Sep 2026 16:09 UTC
9 points
0 comments37 min readLW link
(digitalminds.substack.com)

A non-gen­er­a­tive model as a trusted mon­i­tor for AI Con­trol: Test­ing TypeSafe’s Jev

Venkat T18 Sep 2026 16:06 UTC
15 points
0 comments15 min readLW link

My Reflec­tions Towards the Path to Greatness

Jv Thunder18 Sep 2026 15:33 UTC
0 points
1 comment5 min readLW link

The Prefer­ence Cas­cade Is Only Get­ting Started

Zvi18 Sep 2026 14:40 UTC
66 points
3 comments27 min readLW link
(thezvi.wordpress.com)

An­nounc­ing For­mal Ver­ifi­ca­tion at RESI (The In­sti­tute for Re­spon­si­ble Su­per­in­tel­li­gence)

Adam Chlipala18 Sep 2026 14:32 UTC
13 points
0 comments6 min readLW link

Col­lec­tive Epistemics: Nap­kin Math on In­de­pen­dent Errors

Jonas Hallgren18 Sep 2026 13:42 UTC
26 points
1 comment7 min readLW link

You don’t need a union to go on strike

sudo-nym18 Sep 2026 6:38 UTC
42 points
3 comments1 min readLW link

The Game is Set for a Tar­geted Memetic At­tack on the AI Safety Community

keltan18 Sep 2026 5:43 UTC
139 points
20 comments1 min readLW link

The Align­ment Prob­lem in Align­ment Re­search(ers): a Vol­un­tary­ist Meta-Ethics Perspective

Paul Varkey Parayil18 Sep 2026 4:17 UTC
2 points
0 comments7 min readLW link

Three Hack­ers used Opus 5 to Hack Into OpenAI’s Core Code­base [WSJ]

Linch18 Sep 2026 4:15 UTC
50 points
1 comment1 min readLW link
(www.wsj.com)

The Horse

Character#273618 Sep 2026 2:52 UTC
67 points
4 comments3 min readLW link

Deep re­cur­rent mod­els are less ro­bustly CoT-mon­i­torable than nor­mal CoT mod­els in a toy setting

18 Sep 2026 2:47 UTC
89 points
1 comment11 min readLW link

Ma­chine in­tel­li­gence and the death of hu­man expression

Girard Dorney18 Sep 2026 2:10 UTC
6 points
0 comments6 min readLW link
(extinctiondesk.substack.com)

The Cost of Utopias (a Dia­log)

WillPetillo18 Sep 2026 1:57 UTC
16 points
2 comments17 min readLW link

Two Axes of Align­ment: A Frame­work for Ro­bust Su­per­in­tel­li­gence Alignment

Arihant Gadgade18 Sep 2026 0:59 UTC
7 points
0 comments4 min readLW link

Towards Align­ment Au­dit­ing for RL Environments

Dylan Hawk18 Sep 2026 0:54 UTC
18 points
0 comments12 min readLW link

How to Unclench

Jonny Miller18 Sep 2026 0:52 UTC
13 points
1 comment47 min readLW link
(howtounclench.com)

A web­site to ex­press friend­ship to fu­ture AGI.

FutureFriendofAI18 Sep 2026 0:51 UTC
4 points
0 comments2 min readLW link

Hid­den Knowl­edge? Arrr...

Bobby Faber18 Sep 2026 0:40 UTC
8 points
0 comments2 min readLW link

Su­per­in­tel­li­gence this Christmas

Alexander Gietelink Oldenziel18 Sep 2026 0:06 UTC
40 points
15 comments3 min readLW link

What is (and isn’t) gained by avoid­ing ar­chi­tec­tures with high opaque se­rial depth?

Alek Westover17 Sep 2026 23:52 UTC
11 points
0 comments3 min readLW link

If METR is over­worked, how to alle­vi­ate the bot­tle­neck?

Matthew_Opitz17 Sep 2026 22:56 UTC
17 points
1 comment3 min readLW link

AI is an abun­dance of choice not a 1D spectrum

KatjaGrace17 Sep 2026 22:39 UTC
11 points
0 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

AI can kill us with­out hu­man ex­tinc­tion: P(Catas­tro­phe)

Young Jae Koh17 Sep 2026 22:22 UTC
4 points
0 comments4 min readLW link

Grant­mak­ers aren’t afraid to die

dan.parshall17 Sep 2026 21:52 UTC
23 points
28 comments7 min readLW link
(unsolicitedadvice.ai)

Against AI Risk be­com­ing mainstream

Prometheus17 Sep 2026 21:49 UTC
18 points
5 comments5 min readLW link

YCom­bi­na­tor com­pa­nies still aren’t grow­ing faster due to AI

Xodarap17 Sep 2026 21:31 UTC
13 points
0 comments1 min readLW link

The J-lens offset is the model’s to­ken fre­quency: z-scor­ing helps

Ameya Panchal17 Sep 2026 21:13 UTC
5 points
0 comments10 min readLW link
(ameya-bit.github.io)

A Defense of Grad­ual Disempowerment

Max Harms17 Sep 2026 21:04 UTC
19 points
3 comments6 min readLW link

Swarm Or­ga­ni­za­tion as the Ex­po­nent on Test-Time Compute

Julian Bradshaw17 Sep 2026 20:17 UTC
23 points
1 comment6 min readLW link

“Reg­u­la­tory cap­ture” may be win­ning the Over­ton Window

less_raichu17 Sep 2026 19:56 UTC
−3 points
0 comments3 min readLW link

Good and bad ways to eval­u­ate a defi­ni­tion

Elijah17 Sep 2026 18:17 UTC
22 points
0 comments4 min readLW link

Pac­ing the Fron­tier: A Frame­work & Re­search Agenda

17 Sep 2026 16:36 UTC
40 points
0 comments2 min readLW link
(pacing.tech)

cal­l­congress.ai – the ba­sic ac­tion US res­i­dents can take to help with AI risk

17 Sep 2026 15:26 UTC
66 points
1 comment1 min readLW link
(callcongress.ai)

A Pos­si­ble Solu­tion to the Ob­serv­abil­ity Problem

Isha Yiras Hashem 17 Sep 2026 14:50 UTC
2 points
0 comments7 min readLW link

Emer­gency Me­dia Re­sponse PauseAI Protest Speech

Laiba Rehman ✦ RJ17 Sep 2026 14:42 UTC
12 points
0 comments3 min readLW link

As­tra uses some of its no-CoT ca­pa­bil­ity in practice

StevenW17 Sep 2026 14:21 UTC
9 points
2 comments3 min readLW link

Did Gal­ileo mis­take Saturn’s rings for Jupiter’s Moons?

Alfred Harwood17 Sep 2026 12:56 UTC
37 points
2 comments2 min readLW link

What is it like to be al­ive?

David Balduzzi17 Sep 2026 12:25 UTC
2 points
3 comments12 min readLW link

AI #186: The World Takes Notice

Zvi17 Sep 2026 12:10 UTC
50 points
3 comments55 min readLW link
(thezvi.wordpress.com)

Min­dread­ing is com­ing, what’s the plan?

MathiasKB17 Sep 2026 11:38 UTC
12 points
3 comments1 min readLW link

plz­don­tkil­lus Fel­lows Got ~2M AI Safety Views, Not 21M

Josh Thorsteinson17 Sep 2026 8:17 UTC
41 points
1 comment6 min readLW link

AI as or­derly evac­u­a­tion vs stampede

Richard_Ngo17 Sep 2026 2:40 UTC
161 points
4 comments5 min readLW link
(www.mindthefuture.info)

For Love of the Light­cone, Don’t Par­ti­sanize AI Safety

DanB17 Sep 2026 0:46 UTC
159 points
54 comments14 min readLW link
(unifixion.substack.com)