RSS

The Game is Set for a Tar­geted Memetic At­tack on the AI Safety Community

keltan18 Sep 2026 5:43 UTC
43 points
8 comments1 min readLW link

The Horse

Character#273618 Sep 2026 2:52 UTC
21 points
1 comment3 min readLW link

Deep re­cur­rent mod­els are less ro­bustly CoT-mon­i­torable than nor­mal CoT mod­els in a toy setting

18 Sep 2026 2:47 UTC
73 points
1 comment14 min readLW link

Ma­chine in­tel­li­gence and the death of hu­man expression

Girard Dorney18 Sep 2026 2:10 UTC
6 points
0 comments6 min readLW link
(extinctiondesk.substack.com)

The Cost of Utopias (a Dia­log)

WillPetillo18 Sep 2026 1:57 UTC
17 points
2 comments17 min readLW link

Two Axes of Align­ment: A Frame­work for Ro­bust Su­per­in­tel­li­gence Alignment

Arihant Gadgade18 Sep 2026 0:59 UTC
6 points
0 comments4 min readLW link

Su­per­in­tel­li­gence this Christmas

Alexander Gietelink Oldenziel18 Sep 2026 0:06 UTC
55 points
8 comments3 min readLW link

What is (and isn’t) gained by avoid­ing ar­chi­tec­tures with high opaque se­rial depth?

Alek Westover17 Sep 2026 23:52 UTC
11 points
0 comments3 min readLW link

If METR is over­worked, how to alle­vi­ate the bot­tle­neck?

Matthew_Opitz17 Sep 2026 22:56 UTC
17 points
0 comments3 min readLW link

AI is an abun­dance of choice not a 1D spectrum

KatjaGrace17 Sep 2026 22:39 UTC
9 points
0 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

AI can kill us with­out hu­man ex­tinc­tion: P(Catas­tro­phe)

Young Jae Koh17 Sep 2026 22:22 UTC
3 points
0 comments4 min readLW link

Grant­mak­ers aren’t afraid to die

dan.parshall17 Sep 2026 21:52 UTC
28 points
19 comments5 min readLW link
(unsolicitedadvice.ai)

Against AI Risk be­com­ing mainstream

Prometheus17 Sep 2026 21:49 UTC
16 points
3 comments5 min readLW link

The J-lens offset is the model’s to­ken fre­quency: z-scor­ing helps

Ameya Panchal17 Sep 2026 21:13 UTC
7 points
0 comments10 min readLW link
(ameya-bit.github.io)

A Defense of Grad­ual Disempowerment

Max Harms17 Sep 2026 21:04 UTC
19 points
3 comments6 min readLW link

Swarm Or­ga­ni­za­tion as the Ex­po­nent on Test-Time Compute

Julian Bradshaw17 Sep 2026 20:17 UTC
9 points
1 comment6 min readLW link

Good and bad ways to eval­u­ate a defi­ni­tion

Elijah17 Sep 2026 18:17 UTC
20 points
0 comments4 min readLW link

Pac­ing the Fron­tier: A Frame­work & Re­search Agenda

17 Sep 2026 16:36 UTC
36 points
0 comments2 min readLW link
(pacing.tech)

A Pos­si­ble Solu­tion to the Ob­serv­abil­ity Problem

Isha Yiras Hashem 17 Sep 2026 14:50 UTC
2 points
0 comments7 min readLW link

As­tra uses some of its no-CoT ca­pa­bil­ity in practice

StevenW17 Sep 2026 14:21 UTC
9 points
2 comments3 min readLW link