I want to bitch about the in­creased salience of AI risk

Nathan Young9 Sep 2026 23:14 UTC
17 points
10 comments1 min readLW link

Peo­ple are more wor­ried about AI x-risk than they let on

KatjaGrace9 Sep 2026 22:24 UTC
24 points
1 comment1 min readLW link
(worldspiritsockpuppet.substack.com)

Can a su­per­in­tel­li­gence do THAT?

Eliezer Yudkowsky9 Sep 2026 21:59 UTC
135 points
28 comments13 min readLW link

Fi­nal re­search agenda #3: to­wards au­to­mated Husser­lian CEV

Mitchell_Porter9 Sep 2026 21:52 UTC
13 points
0 comments1 min readLW link

Have de­ranged con­spir­acy the­o­ries ac­tu­ally ex­ploded? The short an­swer: prob­a­bly not.

KatSpartz9 Sep 2026 21:35 UTC
13 points
8 comments4 min readLW link

Con­sti­tu­tional AI Wi­dens Nar­row Se­cret Loy­alty of LLMs

navraj9 Sep 2026 21:05 UTC
7 points
0 comments7 min readLW link

Re­cur­rent KV-cache shar­ing may un­der­mine the bounded-depth ar­gu­ment for CoT monitorability

MBaert9 Sep 2026 21:02 UTC
30 points
2 comments3 min readLW link

The Risks of Anti-AI Vot­ers Turn­ing Against AI Alignment

voikante9 Sep 2026 20:58 UTC
10 points
0 comments48 min readLW link
(voikante.substack.com)

Fight­ing scope and frame con­trol from the fron­tier-labs

Vincent_Bagayoko9 Sep 2026 20:53 UTC
10 points
0 comments9 min readLW link

An Ode to LessWrong

Tamara Sofía Falcone9 Sep 2026 20:47 UTC
15 points
1 comment1 min readLW link

show LW: pro­gram­matic in­ter­net re­search, scry.io

Xyra Sinclair9 Sep 2026 20:44 UTC
24 points
0 comments1 min readLW link

Au­toma­tion and Poli­ti­cal Power

MarkelKori9 Sep 2026 19:55 UTC
10 points
2 comments2 min readLW link

Pre­train­ing an LLM with­out men­tions of consciousness

9 Sep 2026 19:38 UTC
30 points
13 comments7 min readLW link

Bot­tle­necks in AI Safety—Op­por­tu­ni­ties for Impact

Brian Davies9 Sep 2026 17:48 UTC
26 points
1 comment16 min readLW link

Recom­men­da­tions for Peo­ple Get­ting into Tech­ni­cal AI Gover­nance Research

9 Sep 2026 17:16 UTC
36 points
0 comments1 min readLW link

Per­sonal state­ment on join­ing the OpenAI non­profit board

paulfchristiano9 Sep 2026 17:13 UTC
245 points
75 comments2 min readLW link
(paulfchristiano.substack.com)

What’s The Plan for Plan A Leg­is­la­tion?

Yitz9 Sep 2026 16:36 UTC
25 points
4 comments1 min readLW link

Data bot­tle­necks won’t pre­vent an in­tel­li­gence ex­plo­sion (but they will slow it down)

Tom Davidson9 Sep 2026 16:06 UTC
17 points
0 comments20 min readLW link
(www.forethought.org)

“ULC” and AI x-risk

k3nt9 Sep 2026 15:51 UTC
5 points
4 comments5 min readLW link

Fu­tureE­val Spring Re­sults: Pros Beat Bots, but the Gap is Nearly Gone

9 Sep 2026 15:09 UTC
22 points
0 comments21 min readLW link
(www.metaculus.com)

Es­ti­mat­ing GPT-6 As­tra’s no-CoT Time Horizon

9 Sep 2026 13:54 UTC
98 points
7 comments2 min readLW link

Self Hosting

Tomás B.9 Sep 2026 13:45 UTC
145 points
10 comments2 min readLW link

GPT-6 As­tra: The Sys­tem Card, Align­ment and What Comes Next

Zvi9 Sep 2026 13:10 UTC
51 points
3 comments29 min readLW link
(thezvi.wordpress.com)

Bologna Septem­ber Meetup

Luca Petrolati9 Sep 2026 11:35 UTC
1 point
0 comments1 min readLW link

httpi: the in­ter­net pro­to­col to re­duce com­pute from mis­be­hav­ing agents

joao_abrantes9 Sep 2026 7:54 UTC
−15 points
4 comments4 min readLW link

GPT-6 As­tra can do a lot of multi-hop rea­son­ing with­out chain of thought

RohanS9 Sep 2026 5:15 UTC
101 points
5 comments5 min readLW link

One Billion Hemingways

Girard Dorney9 Sep 2026 4:46 UTC
43 points
5 comments8 min readLW link
(extinctiondesk.substack.com)

No, de­tached lin­ear probes won’t save us

RobinHa9 Sep 2026 4:45 UTC
25 points
4 comments3 min readLW link

Chastity Ruth gone, anonymity dropped

Girard Dorney9 Sep 2026 4:39 UTC
4 points
0 comments1 min readLW link

Where to donate on AI if I agree with EY view? (and if I am from Rus­sia)

EniScien9 Sep 2026 2:27 UTC
20 points
9 comments1 min readLW link

Train­ing against the mon­i­tor: What hap­pens dur­ing Obfus­cated Ad­ver­sar­ial Train­ing?

Venkat T9 Sep 2026 1:36 UTC
17 points
0 comments31 min readLW link

How good are slop-ves­ti­ga­tors?

8 Sep 2026 22:13 UTC
140 points
17 comments5 min readLW link

Allow Car­ri­ers on Planes

jefftk8 Sep 2026 21:30 UTC
27 points
0 comments4 min readLW link
(www.jefftk.com)

Is there an ac­cept­able way to store clothes?

KatjaGrace8 Sep 2026 20:43 UTC
3 points
3 comments3 min readLW link
(worldspiritsockpuppet.substack.com)

FLOP Around and Find Out: LLM Train­ing Work­load Size Es­ti­ma­tion With Power Monitoring

William Fowler8 Sep 2026 19:43 UTC
9 points
0 comments19 min readLW link

A Con­cep­tual Frame­work for Rea­son­ing about Ex­plo­ra­tion Hacking

8 Sep 2026 19:13 UTC
38 points
0 comments17 min readLW link

Ex­plo­ra­tion Hack­ing in AI De­bate: Ini­tial Em­pirics and Gen­er­al­i­sa­tion Splitting

8 Sep 2026 19:13 UTC
33 points
0 comments13 min readLW link

Paus­ing AI ASAP is prefer­able to agree­ing to pause at some fu­ture time

Connor Williams8 Sep 2026 18:06 UTC
92 points
8 comments1 min readLW link

At­tune­ment, not alignment

epicurus8 Sep 2026 17:51 UTC
15 points
2 comments8 min readLW link

OpenAI have solved the Navier-Stokes Prob­lem with a sub­stan­tially more pow­er­ful model than As­tra.

fluxxrider8 Sep 2026 17:30 UTC
51 points
16 comments1 min readLW link

Train­ing on probes: Re­search ideas

Charlie Steiner8 Sep 2026 17:06 UTC
27 points
3 comments7 min readLW link

Train­ing on probes: What’s go­ing on

Charlie Steiner8 Sep 2026 17:06 UTC
68 points
0 comments13 min readLW link

The Seven Lan­guages of Comfort

yatharth8 Sep 2026 16:17 UTC
15 points
0 comments7 min readLW link

As­tra and Fable still hack on sim­ple var­i­ants of al­ign­ment evals from 2025

Dean Valentine8 Sep 2026 15:05 UTC
490 points
26 comments2 min readLW link
(goodhartlabs.com)

AI Chased Me From Academia

Mikewins8 Sep 2026 15:01 UTC
39 points
0 comments14 min readLW link

Psy­cholog­i­cal Sup­port for AI Safety Re­searchers Is Ne­glected and Easy to Provide

Ihor Kendiukhov8 Sep 2026 13:19 UTC
83 points
5 comments8 min readLW link

Paper­clips Can’t Fight Back (Why Many Ul­ti­mate Goals Should Lead to Per­va­sive and Di­verse In­tel­li­gence)

Adam Chlipala8 Sep 2026 12:08 UTC
10 points
5 comments8 min readLW link

Now What? Why Field Builders Need More Skin In The Game

Jonas Hallgren8 Sep 2026 11:47 UTC
15 points
0 comments11 min readLW link

As­tra Is Hard to Monitor

Zvi8 Sep 2026 11:30 UTC
47 points
2 comments36 min readLW link
(thezvi.wordpress.com)

Build­ing safe phys­i­cal AI: Open ques­tions, risks, and a call for collaboration

8 Sep 2026 5:48 UTC
4 points
5 comments14 min readLW link
(convergentrobotics.org)