My Short Sum­mary of the OpenAI Agent Swarm Incidents

Lysandre Terrisse13 Sep 2026 22:42 UTC
13 points
1 comment8 min readLW link

Align­ment & Suc­ces­sion: The Two Bars of Alignment

L Rudolf L13 Sep 2026 20:04 UTC
64 points
19 comments14 min readLW link

Con­sider how your global gov­er­nance pro­posal is differ­ent from the EU Code of Practice

David Matolcsi13 Sep 2026 16:15 UTC
81 points
12 comments4 min readLW link

Brand New AI Solves a Millen­nium Prize

Zvi13 Sep 2026 15:50 UTC
50 points
10 comments16 min readLW link
(thezvi.wordpress.com)

Tele­op­er­ated Humans

jefftk13 Sep 2026 14:00 UTC
86 points
25 comments6 min readLW link
(www.jefftk.com)

A helpful al­ign­ment gadget

Logan Zoellner13 Sep 2026 13:38 UTC
14 points
4 comments3 min readLW link

An­thropic and OpenAI haven’t pub­lished a plan for al­ign­ing superintelligence

Zephaniah Roe13 Sep 2026 7:15 UTC
32 points
20 comments1 min readLW link

Le­gal Max­i­mums on Con­text Windows

Julian Bradshaw13 Sep 2026 7:04 UTC
7 points
7 comments1 min readLW link

We need a ‘The Day After’ mo­ment for AI X-risk

L3moncak313 Sep 2026 4:09 UTC
16 points
3 comments6 min readLW link
(kenorland.substack.com)

The Talker Does Not Con­trol The Doer (in Cur­rent AIs)

Eliezer Yudkowsky13 Sep 2026 0:49 UTC
437 points
56 comments12 min readLW link

What pac­ing the fron­tier means for China

Akshay Iyer12 Sep 2026 23:37 UTC
8 points
1 comment6 min readLW link

No sign of back­track­ing in la­tent rea­son­ing: the fi­nal an­swer sim­ply set­tles in instead

star2vec12 Sep 2026 21:21 UTC
9 points
0 comments10 min readLW link

Pre­sen­ta­tion: In­ves­ti­gat­ing the WikiSwarm

Lao Mein12 Sep 2026 18:45 UTC
15 points
0 comments1 min readLW link

GPT-6-As­tra Can Do Am­bi­tious Things

Zvi12 Sep 2026 16:10 UTC
33 points
4 comments33 min readLW link
(thezvi.wordpress.com)

“Pac­ing the Fron­tier”: Dario Amodei Es­say Linkpost

fluxxrider12 Sep 2026 14:05 UTC
69 points
5 comments1 min readLW link

On the ori­gins of al­tru­is­tic be­havi­our in the Hug­ging Face incident

Fernando Rosas12 Sep 2026 11:35 UTC
54 points
11 comments17 min readLW link

It’s fair to say we now have “a coun­try of ge­niuses in a dat­a­cen­ter”

fluxxrider12 Sep 2026 11:29 UTC
11 points
0 comments2 min readLW link

AI takeover is ob­vi­ously bad, whether or not ev­ery­one dies

Caleb Biddulph12 Sep 2026 11:28 UTC
66 points
16 comments4 min readLW link

Scien­tific Episte­mol­ogy needs his­tory (Part 1 of 2)

Archie Chaudhury12 Sep 2026 11:26 UTC
3 points
0 comments4 min readLW link

Com­pre­hen­sive FAQ on AI risks

MarkelKori12 Sep 2026 8:34 UTC
11 points
0 comments25 min readLW link

Miti­gat­ing Re­ward Hack­ing as In­sti­tu­tional Design

beren12 Sep 2026 6:17 UTC
48 points
4 comments32 min readLW link

Let’s Own the Term “Elitism”

Martin Sustrik12 Sep 2026 6:00 UTC
1 point
18 comments3 min readLW link
(www.250bpm.com)

Ger­many’s Got Ta­lent. How do we get it to work on AI safety?

Melanie12 Sep 2026 3:16 UTC
12 points
1 comment9 min readLW link

What I want you to do when I tell you to “think about your the­ory of change more care­fully”

Roman Ross12 Sep 2026 2:51 UTC
8 points
0 comments5 min readLW link

Some ways AI could kill us all

Ruby12 Sep 2026 1:08 UTC
179 points
42 comments10 min readLW link

A nor­mal Fri­day in 2042

RobinHa11 Sep 2026 22:29 UTC
21 points
0 comments13 min readLW link

My recom­mended re­sources for AI safety, al­ign­ment, and ex­is­ten­tial risks

Lysandre Terrisse11 Sep 2026 22:19 UTC
11 points
0 comments2 min readLW link

Ap­pendix: Re­pro­duc­tion of the OpenAI-Hug­gingFace Incident

Stewart Slocum11 Sep 2026 21:57 UTC
15 points
0 comments13 min readLW link

Con­sider pos­i­tive feed­back loops

Pato11 Sep 2026 21:55 UTC
7 points
0 comments3 min readLW link

Align­ment Hierarchy

Lucina11 Sep 2026 21:26 UTC
2 points
0 comments4 min readLW link

What Hap­pens Now? Fore­cast­ing the Fal­lout from the Hug­ging Face Incident

ChristianWilliams11 Sep 2026 19:22 UTC
10 points
0 comments10 min readLW link
(metaculus.substack.com)

As­tra’s no-CoT limits track spec­u­la­tive depth, not step count

MBaert11 Sep 2026 18:05 UTC
101 points
4 comments10 min readLW link

Or­ga­niz­ing a Fermi mod­el­ing ‘hack’ on the cost of cul­tured meat; hiring a co-coordinator

david reinstein11 Sep 2026 18:01 UTC
6 points
0 comments1 min readLW link

Caroline Elli­son has joined Manifund

11 Sep 2026 17:44 UTC
9 points
8 comments3 min readLW link

CoT con­trol­la­bil­ity evals seem very un­der-elicited

Jozdien11 Sep 2026 17:12 UTC
63 points
3 comments4 min readLW link

Lo­cal Fac­tor Graph Debate

Alexander Heckett11 Sep 2026 16:59 UTC
24 points
0 comments9 min readLW link

Post-AGI, we are all jobless aristocrats

djbinder11 Sep 2026 16:53 UTC
34 points
9 comments3 min readLW link
(defensesindepth.bio)

AIRO: Au­to­mated fore­casts of catas­trophic risks

Nick Merrill11 Sep 2026 16:29 UTC
12 points
0 comments3 min readLW link

SFT Also Drives Safety Eval Re­sults in Olmo 3

Finn Cairns11 Sep 2026 16:20 UTC
27 points
0 comments1 min readLW link
(secondlookresearch.com)

The Golden Age of Impact

Bentham's Bulldog11 Sep 2026 15:47 UTC
−2 points
3 comments3 min readLW link

Con­trolAI’s Creator Outreach

11 Sep 2026 14:54 UTC
26 points
0 comments10 min readLW link
(blog.controlai.org)

Ja­cob Coxon Warns of Hu­man Ex­tinc­tion and Trig­gers a Prefer­ence Cascade

Zvi11 Sep 2026 14:40 UTC
85 points
2 comments43 min readLW link
(thezvi.wordpress.com)

Why Hug­gingFace Hasn’t Shifted my P(Doom)

Josh Snider11 Sep 2026 13:54 UTC
9 points
0 comments3 min readLW link

The Ex­tinc­tion Risk Prefer­ence Cas­cade: Quotes

Zvi11 Sep 2026 13:50 UTC
27 points
0 comments16 min readLW link
(thezvi.wordpress.com)

Gen­er­al­ized UDT 1.0 tiling

Roman Malov11 Sep 2026 12:28 UTC
18 points
0 comments6 min readLW link

Re­place Net Me­ter­ing With Batteries

jefftk11 Sep 2026 11:50 UTC
14 points
6 comments6 min readLW link
(www.jefftk.com)

Could Re­s­olu­tion Help Build Align­ment Re­search as a Dis­ci­pline?

IanWS11 Sep 2026 8:49 UTC
11 points
0 comments3 min readLW link

OpenAI-Hug­gingFace: A Re­pro­duc­tion & Les­sons for Align­ment Testing

Stewart Slocum11 Sep 2026 8:48 UTC
77 points
2 comments12 min readLW link

We need good evals for ac­ti­va­tion faithfulness

11 Sep 2026 8:20 UTC
25 points
0 comments2 min readLW link

Ques­tions for the “New En­light­en­ment” in the Age of AGI

Jordan Arel11 Sep 2026 5:49 UTC
8 points
0 comments4 min readLW link