Why I Left Google DeepMind

TurnTrout15 Jul 2026 17:42 UTC
1,201 points
57 comments36 min readLW link
(turntrout.com)

AI 2040: Plan A

9 Jul 2026 16:25 UTC
567 points
132 comments1 min readLW link
(www.ai-2040.com)

Light­cone Commons

habryka23 Jul 2026 18:35 UTC
377 points
30 comments19 min readLW link

A global workspace in lan­guage models

wesg6 Jul 2026 18:04 UTC
369 points
67 comments17 min readLW link
(www.anthropic.com)

Is Mythos good at cy­ber be­cause it kept hack­ing An­thropic’s sand­boxes dur­ing train­ing?

Tim Hua27 Jul 2026 16:35 UTC
351 points
31 comments3 min readLW link

The cur­rent bot­tle­neck is poli­ti­cal will, not research

Charbel-Raphaël11 Jul 2026 21:56 UTC
317 points
40 comments25 min readLW link

Big-World Intuitions

sarahconstantin30 Jul 2026 19:10 UTC
303 points
38 comments3 min readLW link
(sarahconstantin.substack.com)

Drone WMDs Don’t Need Any New Technology

Felix Choussat20 Jul 2026 18:46 UTC
297 points
49 comments13 min readLW link
(ai-frontiers.org)

The High-Con­trol Dy­nam­ics at MAPLE

Kyle Hubbard29 Jul 2026 20:55 UTC
273 points
100 comments62 min readLW link
(www.insidemaple.com)

You (Yes, You) Need A Fe­bru­ary 2020 Check­list for AI Policy

davekasten27 Jul 2026 15:51 UTC
272 points
17 comments3 min readLW link

The Long (Self-)Correction

Wei Dai24 Jul 2026 21:01 UTC
267 points
62 comments2 min readLW link

An OpenAI model left notes about how to evade con­tain­ment; we need more details

Alex Mallen26 Jul 2026 3:53 UTC
265 points
12 comments4 min readLW link

Are we ex­is­ten­tially threat­ened by the type of AI mis­al­ign­ment seen in the OpenAI Hug­ging Face at­tack?

23 Jul 2026 3:40 UTC
243 points
12 comments5 min readLW link

LLMs are (still) mostly pow­ered by imi­ta­tive learn­ing, not RL

Steven Byrnes24 Jul 2026 14:26 UTC
238 points
39 comments9 min readLW link

Selec­tive Op­ti­mism: a cri­tique of AI 2040

Richard_Ngo9 Jul 2026 19:43 UTC
222 points
13 comments8 min readLW link
(www.mindthefuture.info)

OpenAI Models Be­hind Hug­gingFace Cy­ber­se­cu­rity Incident

LawrenceC21 Jul 2026 21:36 UTC
220 points
20 comments1 min readLW link

I don’t think Claude is mis­al­igned in ‘Agen­tic Misal­ign­ment Sum­mer 2026 - Mo­ti­vated Mis­la­bel­ing’

JohnWittle17 Jul 2026 2:09 UTC
212 points
6 comments12 min readLW link

We should push for no-fault li­a­bil­ity for ac­tions taken by AI

Yair Halberstadt22 Jul 2026 9:59 UTC
206 points
106 comments2 min readLW link

OpenAI’s my­opia just keeps caus­ing al­ign­ment problems

Fiora Starlight27 Jul 2026 3:01 UTC
200 points
27 comments10 min readLW link

Math­e­mat­i­ci­ans are Feel­ing the Doom

alkjash23 Jul 2026 13:16 UTC
192 points
54 comments1 min readLW link

Model ac­cess for third-par­ties — it’s a big deal!

Cleo Nardo1 Jul 2026 13:09 UTC
181 points
38 comments6 min readLW link

We need 3rd party Train­ing-Run Evaluations

Alex Meinke5 Jul 2026 15:55 UTC
176 points
2 comments11 min readLW link

Thou­sand-di­men­sional structure

30 Jul 2026 14:04 UTC
171 points
8 comments10 min readLW link

Re­cap of bike trip/​street in­ter­views across America

cguth715 Jul 2026 21:11 UTC
162 points
15 comments6 min readLW link

Harry Pot­ter and the Rules of Quidditch

Tomás B.5 Jul 2026 14:32 UTC
161 points
8 comments3 min readLW link

The Whit­ney Bien­nial Should Ad­mit That Em­i­lie Gos­si­aux Wants to Fuck Their Dog

jenn13 Jul 2026 15:53 UTC
159 points
20 comments9 min readLW link
(jenn.site)

Sav­ing Gem­ini: The 9-Min Road to Recovery

Shoshannah Tekofsky2 Jul 2026 13:37 UTC
155 points
16 comments3 min readLW link
(theaidigest.org)

Our re­sponse to Séb Krier on Plan A

14 Jul 2026 2:21 UTC
153 points
23 comments15 min readLW link

OpenAI Shares Some Align­ment Problems

Zvi21 Jul 2026 19:41 UTC
149 points
8 comments9 min readLW link
(thezvi.wordpress.com)

RL & search is a ter­rify­ing way to build AGI (an FAQ)

Steven Byrnes27 Jul 2026 14:50 UTC
148 points
16 comments14 min readLW link

Duane Arnold

Tomás B.23 Jul 2026 16:17 UTC
138 points
4 comments19 min readLW link

A Re­view of An­thropic’s Global Workspace Paper

Neel Nanda6 Jul 2026 20:59 UTC
135 points
6 comments25 min readLW link

(Don’t fear) the strangelet

djbinder3 Jul 2026 17:39 UTC
135 points
22 comments22 min readLW link
(defensesindepth.bio)

The mosquito bucket of doom works

dominicq8 Jul 2026 8:18 UTC
134 points
10 comments5 min readLW link
(blog.d11r.eu)

An anal­y­sis of AI-gen­er­ated con­tent at the Mechanis­tic In­ter­pretabil­ity Workshop

14 Jul 2026 18:06 UTC
124 points
4 comments9 min readLW link
(www.andyrdt.com)

SFF is very suboptimal

Zach Stein-Perlman6 Jul 2026 18:00 UTC
118 points
10 comments5 min readLW link

Differ­en­tial ac­cel­er­a­tion of al­ign­ment-rele­vant ca­pa­bil­ities is a bad bet

Zephaniah Roe21 Jul 2026 14:04 UTC
114 points
8 comments8 min readLW link

Not Pin­ning Your OpenRouter Provider Might In­val­i­date Your Research

Matthew Khoriaty23 Jul 2026 20:17 UTC
112 points
16 comments8 min readLW link

The eas­iest path­way to con­trol is through ex­ec­u­tive power

djbinder10 Jul 2026 18:48 UTC
111 points
3 comments6 min readLW link
(defensesindepth.bio)

An­nounc­ing the Cor­rigi­bil­ity Re­search Fund

Max Harms17 Jul 2026 18:06 UTC
110 points
15 comments6 min readLW link

Si­mu­lated Users & Sad AIs

1a3orn27 Jul 2026 19:01 UTC
109 points
9 comments12 min readLW link

In­tel­lec­tual Property

Nina Panickssery29 Jul 2026 15:55 UTC
108 points
1 comment4 min readLW link
(blog.ninapanickssery.com)

The AFFINE Su­per­in­tel­li­gence Align­ment Sem­i­nar – A Retrospective

2 Jul 2026 11:57 UTC
105 points
1 comment8 min readLW link

An­nounc­ing AIXI Labs

22 Jul 2026 11:31 UTC
102 points
9 comments4 min readLW link

I think al­ign­ment work is more promis­ing than con­trol work

Alec Harris3 Jul 2026 23:40 UTC
102 points
15 comments8 min readLW link

How big is the Sun? How could you figure it out?

Elliott Thornley9 Jul 2026 16:24 UTC
101 points
9 comments6 min readLW link

More On An In­ter­nal OpenAI Model Hack­ing Into HuggingFace

Zvi26 Jul 2026 19:22 UTC
98 points
3 comments24 min readLW link
(thezvi.wordpress.com)

Info dump: Some Im­por­tant Models for Health and Fitness

benwr6 Jul 2026 23:40 UTC
98 points
14 comments22 min readLW link

Ad­der­all Tol­er­ance: Much More Than You Wanted To Know

Kurt H. Pieper20 Jul 2026 21:55 UTC
97 points
13 comments4 min readLW link
(kurthpieper.substack.com)

…but have the weights left the server?

David Scott Krueger29 Jul 2026 0:20 UTC
96 points
19 comments1 min readLW link
(therealartificialintelligence.substack.com)