Why I Left Google DeepMind

TurnTrout15 Jul 2026 17:42 UTC
1,200 points
57 comments36 min readLW link
(turntrout.com)

My jour­ney to the microwave al­ter­nate timeline

Malmesbury10 Feb 2026 17:59 UTC
795 points
58 comments10 min readLW link

Cur­rent AIs seem pretty mis­al­igned to me

ryan_greenblatt15 Apr 2026 15:14 UTC
768 points
83 comments27 min readLW link

How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
758 points
102 comments12 min readLW link

How Go Play­ers Disem­power Them­selves to AI

Ashe Vazquez Nuñez1 May 2026 23:24 UTC
746 points
79 comments8 min readLW link

Only Law Can Prevent Extinction

Eliezer Yudkowsky13 Apr 2026 20:57 UTC
744 points
110 comments20 min readLW link

What just hap­pened? A ret­ro­spec­tive of AI alignment

Richard_Ngo9 Aug 2026 15:58 UTC
635 points
152 comments16 min readLW link

The Terrarium

Caleb Biddulph26 Mar 2026 18:08 UTC
611 points
54 comments21 min readLW link

Brief in­de­pen­dent in­ves­ti­ga­tion of agents’ be­hav­ior, rea­son­ing and col­lab­o­ra­tion in the OpenAI /​ Hug­ging Face hack­ing incident

26 Aug 2026 19:40 UTC
589 points
66 comments3 min readLW link
(metr.org)

Here’s to the Polypropy­lene Makers

jefftk27 Feb 2026 4:00 UTC
566 points
19 comments2 min readLW link
(www.jefftk.com)

AI 2040: Plan A

9 Jul 2026 16:25 UTC
562 points
132 comments1 min readLW link
(www.ai-2040.com)

Less Dead

Aurelia11 Mar 2026 5:07 UTC
555 points
140 comments8 min readLW link

Ar­gu­ments for P

Cleo Nardo5 Aug 2026 19:00 UTC
548 points
66 comments2 min readLW link

Sur­pris­ing facts about the slave trade

Joseph Miller26 Jun 2026 1:34 UTC
466 points
52 comments8 min readLW link

What just hap­pened? Prag­ma­tism and Pessimization

Richard_Ngo24 Aug 2026 2:03 UTC
466 points
176 comments27 min readLW link

What I did in the he­do­nium shock­wave, by Emma, age six and a half

ozymandias13 Apr 2026 16:47 UTC
457 points
48 comments5 min readLW link
(ozybrennan.substack.com)

Do not con­quer what you can­not defend

habryka16 Apr 2026 4:13 UTC
440 points
73 comments6 min readLW link

As­tra and Fable still hack on sim­ple var­i­ants of al­ign­ment evals from 2025

Dean Valentine8 Sep 2026 15:05 UTC
437 points
19 comments2 min readLW link
(goodhartlabs.com)

AI pause: the case for ASAP

KatjaGrace24 Jun 2026 21:31 UTC
418 points
55 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

Did Claude 3 Opus al­ign it­self via gra­di­ent hack­ing?

Fiora Starlight21 Feb 2026 22:24 UTC
416 points
49 comments20 min readLW link

Fore­cast­ing is Way Over­rated, and We Should Stop Fund­ing It

mabramov25 Apr 2026 22:39 UTC
412 points
99 comments5 min readLW link

A Mechanis­tic Ex­pla­na­tion of Prompt In­jec­tion (and why you should study roles)

22 Jun 2026 14:09 UTC
403 points
62 comments16 min readLW link

The Owned Ones

Eliezer Yudkowsky12 May 2026 17:56 UTC
399 points
53 comments6 min readLW link

How AI Is Learn­ing to Think in Secret

Nicholas Andresen6 Jan 2026 16:31 UTC
395 points
33 comments18 min readLW link
(nickandresen.substack.com)

Re­turn­ing to ARC

paulfchristiano4 Aug 2026 22:27 UTC
387 points
36 comments9 min readLW link

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
382 points
31 comments6 min readLW link

Women should be able to open things

KatjaGrace21 May 2026 3:50 UTC
379 points
142 comments2 min readLW link
(worldspiritsockpuppet.com)

Light­cone Commons

habryka23 Jul 2026 18:35 UTC
377 points
30 comments19 min readLW link

In My Misan­thropy Era

jenn4 Jan 2026 18:34 UTC
376 points
157 comments8 min readLW link
(jenn.site)

Ir­re­triev­abil­ity; or, Mur­phy’s Curse of Oneshot­ness upon ASI

Eliezer Yudkowsky4 May 2026 22:11 UTC
370 points
132 comments22 min readLW link

A global workspace in lan­guage models

wesg6 Jul 2026 18:04 UTC
369 points
67 comments17 min readLW link
(www.anthropic.com)

Mnemonic por­traits for 19,023 hu­man genes

Brinedew28 May 2026 22:16 UTC
363 points
29 comments15 min readLW link

AI found 12 of 12 OpenSSL zero-days (while curl can­cel­led its bug bounty)

Stanislav Fort27 Jan 2026 20:21 UTC
359 points
25 comments8 min readLW link

llm as­sis­tant per­sonas seem in­creas­ingly in­co­her­ent (some sub­jec­tive ob­ser­va­tions)

nostalgebraist29 Apr 2026 3:53 UTC
355 points
83 comments9 min readLW link

Is Mythos good at cy­ber be­cause it kept hack­ing An­thropic’s sand­boxes dur­ing train­ing?

Tim Hua27 Jul 2026 16:35 UTC
351 points
31 comments3 min readLW link

It’s nice of you to worry about me, but I re­ally do have a life

Viliam4 May 2026 21:14 UTC
345 points
66 comments4 min readLW link

On The In­de­pen­dence Axiom

Ihor Kendiukhov8 Mar 2026 14:38 UTC
336 points
93 comments23 min readLW link

Canada Lost Its Measles Elimi­na­tion Sta­tus Be­cause We Don’t Have Enough Nurses Who Speak Low German

jenn25 Jan 2026 18:33 UTC
335 points
24 comments7 min readLW link
(www.jenn.site)

Morale

J Bostock12 Apr 2026 20:15 UTC
334 points
47 comments2 min readLW link

Trees are mostly made of air and a gen­er­al­iz­able les­son for AI safety

Zephaniah Roe29 May 2026 4:08 UTC
324 points
62 comments4 min readLW link

An­noy­ingly Prin­ci­pled Peo­ple, and what be­falls them

Raemon13 Apr 2026 17:35 UTC
322 points
75 comments5 min readLW link

My hobby: run­ning de­ranged surveys

leogao27 Mar 2026 0:41 UTC
321 points
66 comments9 min readLW link

Cat-Bel­ling Problems

Eliezer Yudkowsky3 Sep 2026 21:20 UTC
320 points
92 comments21 min readLW link

The Talker Does Not Con­trol The Doer (in Cur­rent AIs)

Eliezer Yudkowsky13 Sep 2026 0:49 UTC
320 points
43 comments12 min readLW link

The cur­rent bot­tle­neck is poli­ti­cal will, not research

Charbel-Raphaël11 Jul 2026 21:56 UTC
317 points
40 comments25 min readLW link

Models find­ing soft­ware vuln­er­a­bil­ities is not the pri­mary source of cy­ber­se­cu­rity risk

Dean Valentine14 May 2026 3:39 UTC
316 points
25 comments2 min readLW link

Ada Palmer: In­vent­ing the Renaissance

Martin Sustrik26 Jan 2026 4:40 UTC
316 points
22 comments13 min readLW link
(www.250bpm.com)

Re­s­olu­tion (fka Se­quent): scale and au­toma­tion for higher con­fi­dence in alignment

10 Jun 2026 15:37 UTC
310 points
6 comments11 min readLW link
(sequent.org)

Why I think polyamory is net nega­tive for most peo­ple who try it

KatSpartz30 Aug 2026 6:40 UTC
305 points
67 comments8 min readLW link

Day­care illnesses

Nina Panickssery13 Apr 2026 0:46 UTC
304 points
40 comments4 min readLW link
(blog.ninapanickssery.com)