As­tra and Fable still hack on sim­ple var­i­ants of al­ign­ment evals from 2025

Dean Valentine8 Sep 2026 15:05 UTC
437 points
19 comments2 min readLW link
(goodhartlabs.com)

The Talker Does Not Con­trol The Doer (in Cur­rent AIs)

Eliezer Yudkowsky13 Sep 2026 0:49 UTC
324 points
43 comments12 min readLW link

Cat-Bel­ling Problems

Eliezer Yudkowsky3 Sep 2026 21:20 UTC
320 points
92 comments21 min readLW link

Dis­cov­ery Of A New OpenAI Agent Mes­sage Board

Capybasilisk4 Sep 2026 14:46 UTC
289 points
39 comments1 min readLW link
(collusion.wiki)

PauseAI Has ‘offi­cially dis­endorsed’ PauseAI-US

nem1 Sep 2026 15:51 UTC
277 points
215 comments8 min readLW link

Re­s­olu­tion has a new Agent Foun­da­tions team

Jeremy Gillen3 Sep 2026 0:53 UTC
244 points
4 comments2 min readLW link
(resolution.org)

Per­sonal state­ment on join­ing the OpenAI non­profit board

paulfchristiano9 Sep 2026 17:13 UTC
243 points
74 comments2 min readLW link
(paulfchristiano.substack.com)

Sen. Bernie San­ders (I-VT) and Rep. Greg Casar (D-TX) in­tro­duce leg­is­la­tion to ban Ar­tifi­cial Su­per­in­tel­li­gence and tem­porar­ily pause ad­vanced AI de­vel­op­ment

Matrice Jacobine3 Sep 2026 14:41 UTC
225 points
31 comments2 min readLW link
(www.sanders.senate.gov)

As­tra can do a con­cern­ing amount with no chain of thought

Neel Nanda10 Sep 2026 2:31 UTC
203 points
19 comments8 min readLW link

Evaluation

Nina Panickssery5 Sep 2026 16:25 UTC
185 points
7 comments3 min readLW link
(blog.ninapanickssery.com)

Doom as a bad method not a utopia trade-off

KatjaGrace10 Sep 2026 20:14 UTC
177 points
21 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

How con­cerned should we be about As­tra’s re­cur­rent ar­chi­tec­ture?

Rauno Arike2 Sep 2026 18:28 UTC
168 points
28 comments10 min readLW link

The Lo­cally Op­ti­mal Dis­cur­sive Posture

deanball10 Sep 2026 23:01 UTC
159 points
32 comments19 min readLW link

Let’s talk about the AI co­or­di­na­tion problem

KatjaGrace4 Sep 2026 20:16 UTC
159 points
12 comments2 min readLW link

Ex­plain­ing Knigh­ti­anism on one foot

Richard_Ngo2 Sep 2026 0:53 UTC
159 points
23 comments11 min readLW link
(www.mindthefuture.info)

Steer­ing to­wards “au­to­mated grad­ing” de­grades alignment

3 Sep 2026 18:17 UTC
154 points
34 comments6 min readLW link

Some ways AI could kill us all

Ruby12 Sep 2026 1:08 UTC
151 points
29 comments10 min readLW link

Train­ing a Misal­igned Re­ward Seeker

1 Sep 2026 1:41 UTC
149 points
11 comments2 min readLW link
(alignment.anthropic.com)

The Scram­ble: get­ting in po­si­tion to pace the frontier

Peter Wildeford7 Sep 2026 20:23 UTC
142 points
9 comments10 min readLW link

Self Hosting

Tomás B.9 Sep 2026 13:45 UTC
141 points
10 comments2 min readLW link

How good are slop-ves­ti­ga­tors?

8 Sep 2026 22:13 UTC
140 points
17 comments5 min readLW link

Pro­posal for track­ing the effects of ar­chi­tec­ture on monitorability

10 Sep 2026 17:18 UTC
138 points
4 comments3 min readLW link
(www.redwoodresearch.org)

As­tra is much bet­ter at rea­son­ing with filler to­kens than pre­vi­ous models

10 Sep 2026 22:21 UTC
132 points
6 comments2 min readLW link

Can a su­per­in­tel­li­gence do THAT?

Eliezer Yudkowsky9 Sep 2026 21:59 UTC
129 points
24 comments13 min readLW link

Dear God, Please Do Not Re­sign In Protest

Kabir Kumar7 Sep 2026 21:55 UTC
121 points
76 comments2 min readLW link

OpenAI and the Wiki Incident

Zvi6 Sep 2026 20:02 UTC
111 points
9 comments16 min readLW link
(thezvi.wordpress.com)

GPT-6 As­tra can do a lot of multi-hop rea­son­ing with­out chain of thought

RohanS9 Sep 2026 5:15 UTC
101 points
5 comments5 min readLW link

Should safety re­searchers quit fron­tier labs re. warn­ing shots?

Ryan Kidd5 Sep 2026 1:17 UTC
100 points
59 comments4 min readLW link

As­tra’s no-CoT limits track spec­u­la­tive depth, not step count

MBaert11 Sep 2026 18:05 UTC
98 points
4 comments10 min readLW link

Paus­ing AI ASAP is prefer­able to agree­ing to pause at some fu­ture time

Connor Williams8 Sep 2026 18:06 UTC
98 points
8 comments1 min readLW link

Es­ti­mat­ing GPT-6 As­tra’s no-CoT Time Horizon

9 Sep 2026 13:54 UTC
98 points
7 comments2 min readLW link

Salad days

Zephaniah Roe1 Sep 2026 5:26 UTC
96 points
6 comments3 min readLW link

The Align­ment Jour­nal: Or­ga­ni­za­tion, Per­son­nel, and Scope

1 Sep 2026 20:49 UTC
95 points
5 comments11 min readLW link
(blog.alignmentjournal.org)

Kairos has raised $50M to build tal­ent in­fras­truc­ture for AI safety (and we’re hiring!)

agucova2 Sep 2026 20:13 UTC
94 points
8 comments8 min readLW link

Heat Dis­si­pa­tion Is the Main Con­straint in In­ter­stel­lar Travel

Pasha Kamyshev6 Sep 2026 23:03 UTC
92 points
28 comments7 min readLW link

Psy­cholog­i­cal Sup­port for AI Safety Re­searchers Is Ne­glected and Easy to Provide

Ihor Kendiukhov8 Sep 2026 13:19 UTC
82 points
5 comments8 min readLW link

Ja­cob Coxon Warns of Hu­man Ex­tinc­tion and Trig­gers a Prefer­ence Cascade

Zvi11 Sep 2026 14:40 UTC
80 points
2 comments43 min readLW link
(thezvi.wordpress.com)

The Mag­nus Challenge

Taylor G. Lunt7 Sep 2026 19:02 UTC
78 points
12 comments6 min readLW link

An op­er­a­tional­iza­tion of opaque se­rial depth

10 Sep 2026 17:26 UTC
73 points
1 comment12 min readLW link
(www.redwoodresearch.org)

The Geom­e­try of Non­er­godic Composition

10 Sep 2026 18:12 UTC
73 points
0 comments15 min readLW link
(simplex.pub)

A pro­posal for a highly effec­tive AI safety org

ceselder2 Sep 2026 21:34 UTC
72 points
14 comments2 min readLW link

“An Alien Mind” from OAI chief sci­en­tist seems newly cau­tious on alignment

Seth Herd7 Sep 2026 4:13 UTC
72 points
13 comments2 min readLW link
(openai.com)

“Pac­ing the Fron­tier”: Dario Amodei Es­say Linkpost

fluxxrider12 Sep 2026 14:05 UTC
69 points
5 comments1 min readLW link

Don’t be the vi­tamin B guy

HedonicEscalator2 Sep 2026 2:02 UTC
68 points
6 comments6 min readLW link
(hedonicescalator.substack.com)

Train­ing on probes: What’s go­ing on

Charlie Steiner8 Sep 2026 17:06 UTC
67 points
0 comments13 min readLW link

Cat­e­gor­i­cal taboos are much bet­ter than thresh­old taboos: neu­ralese edition

Linch10 Sep 2026 18:57 UTC
66 points
0 comments4 min readLW link

Al­most no­body is funded to figure out what work would solve alignment

Seth Herd4 Sep 2026 15:57 UTC
64 points
20 comments4 min readLW link

AI takeover is ob­vi­ously bad, whether or not ev­ery­one dies

Caleb Biddulph12 Sep 2026 11:28 UTC
63 points
16 comments4 min readLW link

CoT con­trol­la­bil­ity evals seem very un­der-elicited

Jozdien11 Sep 2026 17:12 UTC
63 points
3 comments4 min readLW link

A Year of Atheism

Laiba Rehman ✦ RJ1 Sep 2026 16:19 UTC
62 points
4 comments8 min readLW link