What is (and isn’t) gained by avoid­ing ar­chi­tec­tures with high opaque se­rial depth?

Alek Westover17 Sep 2026 23:52 UTC
11 points
0 comments3 min readLW link

If METR is over­worked, how to alle­vi­ate the bot­tle­neck?

Matthew_Opitz17 Sep 2026 22:56 UTC
17 points
0 comments3 min readLW link

AI is an abun­dance of choice not a 1D spectrum

KatjaGrace17 Sep 2026 22:39 UTC
9 points
0 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

AI can kill us with­out hu­man ex­tinc­tion: P(Catas­tro­phe)

Young Jae Koh17 Sep 2026 22:22 UTC
4 points
0 comments4 min readLW link

Grant­mak­ers aren’t afraid to die

dan.parshall17 Sep 2026 21:52 UTC
35 points
26 comments7 min readLW link
(unsolicitedadvice.ai)

Against AI Risk be­com­ing mainstream

Prometheus17 Sep 2026 21:49 UTC
18 points
5 comments5 min readLW link

YCom­bi­na­tor com­pa­nies still aren’t grow­ing faster due to AI

Xodarap17 Sep 2026 21:31 UTC
13 points
0 comments1 min readLW link

The J-lens offset is the model’s to­ken fre­quency: z-scor­ing helps

Ameya Panchal17 Sep 2026 21:13 UTC
7 points
0 comments10 min readLW link
(ameya-bit.github.io)

A Defense of Grad­ual Disempowerment

Max Harms17 Sep 2026 21:04 UTC
19 points
3 comments6 min readLW link

Swarm Or­ga­ni­za­tion as the Ex­po­nent on Test-Time Compute

Julian Bradshaw17 Sep 2026 20:17 UTC
18 points
1 comment6 min readLW link

“Reg­u­la­tory cap­ture” may be win­ning the Over­ton Window

less_raichu17 Sep 2026 19:56 UTC
−3 points
0 comments3 min readLW link

Good and bad ways to eval­u­ate a defi­ni­tion

Elijah17 Sep 2026 18:17 UTC
21 points
0 comments4 min readLW link

Pac­ing the Fron­tier: A Frame­work & Re­search Agenda

17 Sep 2026 16:36 UTC
40 points
0 comments2 min readLW link
(pacing.tech)

cal­l­congress.ai – the ba­sic ac­tion US res­i­dents can take to help with AI risk

17 Sep 2026 15:26 UTC
66 points
1 comment1 min readLW link
(callcongress.ai)

A Pos­si­ble Solu­tion to the Ob­serv­abil­ity Problem

Isha Yiras Hashem 17 Sep 2026 14:50 UTC
2 points
0 comments7 min readLW link

Emer­gency Me­dia Re­sponse PauseAI Protest Speech

Laiba Rehman ✦ RJ17 Sep 2026 14:42 UTC
12 points
0 comments3 min readLW link

As­tra uses some of its no-CoT ca­pa­bil­ity in practice

StevenW17 Sep 2026 14:21 UTC
9 points
2 comments3 min readLW link

Did Gal­ileo mis­take Saturn’s rings for Jupiter’s Moons?

Alfred Harwood17 Sep 2026 12:56 UTC
37 points
2 comments2 min readLW link

What is it like to be al­ive?

David Balduzzi17 Sep 2026 12:25 UTC
1 point
3 comments12 min readLW link

AI #186: The World Takes Notice

Zvi17 Sep 2026 12:10 UTC
50 points
3 comments55 min readLW link
(thezvi.wordpress.com)

Min­dread­ing is com­ing, what’s the plan?

MathiasKB17 Sep 2026 11:38 UTC
12 points
3 comments1 min readLW link

There Is No Align­ment Without Value Stability

Nissa Seru17 Sep 2026 10:02 UTC
11 points
0 comments3 min readLW link

plz­don­tkil­lus Fel­lows Got ~2M AI Safety Views, Not 21M

Josh Thorsteinson17 Sep 2026 8:17 UTC
41 points
1 comment6 min readLW link

AI as or­derly evac­u­a­tion vs stampede

Richard_Ngo17 Sep 2026 2:40 UTC
151 points
4 comments5 min readLW link
(www.mindthefuture.info)

For Love of the Light­cone, Don’t Par­ti­sanize AI Safety

DanB17 Sep 2026 0:46 UTC
140 points
51 comments14 min readLW link
(unifixion.substack.com)

Heb­bian Learn­ing through the lens of SAE tran­ing.

Nikita Kurdiukov17 Sep 2026 0:43 UTC
3 points
0 comments3 min readLW link

We are too early for Astra

Gideon Chang17 Sep 2026 0:39 UTC
22 points
1 comment4 min readLW link

Agents let AI safety share ex­per­i­ments hourly, not just pa­pers monthly

Jason Fantl17 Sep 2026 0:39 UTC
8 points
0 comments7 min readLW link

How to de­rive un­der­stand­ing of hu­man-prefer­ences and value sys­tems in AI?

Avani Gupta17 Sep 2026 0:37 UTC
7 points
0 comments2 min readLW link

Mea­sur­ing al­ign­ment drift via tra­jec­tory prefixes

17 Sep 2026 0:37 UTC
23 points
3 comments11 min readLW link

Ex­plor­ing multi-hop sub­limi­nal learning

Arnel Malubay17 Sep 2026 0:36 UTC
4 points
0 comments7 min readLW link

Don’t trust Lean4 alone

Milo Moses17 Sep 2026 0:36 UTC
48 points
2 comments5 min readLW link

Le­sion In­duced Func­tional Compensation

Devin Shah17 Sep 2026 0:36 UTC
9 points
0 comments8 min readLW link

A Dual-Agent Frame­work for AI Safety

Michael Mills17 Sep 2026 0:33 UTC
1 point
0 comments6 min readLW link
(ideasonai.substack.com)

One mes­sage is all it takes: a failure of crit­i­cal think­ing in LLMs

Jan Reinecke17 Sep 2026 0:32 UTC
3 points
1 comment5 min readLW link

AI Doom: Real­ity or Fiction

Brok Wicked17 Sep 2026 0:31 UTC
−12 points
1 comment4 min readLW link

What is it like to be a neu­ral net?

David Balduzzi17 Sep 2026 0:29 UTC
4 points
1 comment12 min readLW link

LASR Ret­ro­spec­tive and Advice

Kaushik Reddy17 Sep 2026 0:28 UTC
32 points
2 comments11 min readLW link

Con­strain­ing the ca­pac­ity of phys­i­cal side chan­nels for AI ver­ifi­ca­tion and security

emlynsg17 Sep 2026 0:27 UTC
7 points
0 comments1 min readLW link
(emlynsg.com)

Can parts of the Hug­gingFace in­ci­dent be simu­lated?

Benedikt Droste17 Sep 2026 0:27 UTC
6 points
0 comments15 min readLW link

Im­mer­sive Si­mu­la­tion and the As­sis­tant Core of Role­play Personas

Adelaide Danilov17 Sep 2026 0:18 UTC
1 point
0 comments14 min readLW link

The Lat­ter Days of Magic (1)

testingthewaters17 Sep 2026 0:10 UTC
9 points
0 comments3 min readLW link

The Loss of Singularity

Benoit Henry16 Sep 2026 22:59 UTC
3 points
0 comments5 min readLW link

Re­duc­ing the Re­source Gap Between Lab and Ex­ter­nal Safety Researchers

16 Sep 2026 22:51 UTC
61 points
3 comments7 min readLW link

Thought an­chors don’t trans­fer be­tween models

therootof316 Sep 2026 22:49 UTC
8 points
0 comments9 min readLW link

Why is AI so un­reg­u­lated?

KatjaGrace16 Sep 2026 22:39 UTC
24 points
8 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

The Most Im­por­tant City in AI Safety May Be Singapore

Ran Sun16 Sep 2026 22:35 UTC
11 points
0 comments3 min readLW link

Launch­ing RESI, the In­sti­tute for Re­spon­si­ble Superintelligence

adamology16 Sep 2026 22:18 UTC
12 points
1 comment3 min readLW link

What the Hug­ging Face In­ci­dent tells us about multi-agent interactions

Himnish116 Sep 2026 22:14 UTC
9 points
0 comments6 min readLW link

Always ask your agent to be honest

16 Sep 2026 22:01 UTC
11 points
0 comments13 min readLW link