How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
767 points
103 comments12 min readLW link

What just hap­pened? A ret­ro­spec­tive of AI alignment

Richard_Ngo9 Aug 2026 15:58 UTC
640 points
152 comments16 min readLW link

Brief in­de­pen­dent in­ves­ti­ga­tion of agents’ be­hav­ior, rea­son­ing and col­lab­o­ra­tion in the OpenAI /​ Hug­ging Face hack­ing incident

26 Aug 2026 19:40 UTC
591 points
67 comments3 min readLW link
(metr.org)

Ar­gu­ments for P

Cleo Nardo5 Aug 2026 19:00 UTC
548 points
66 comments2 min readLW link

What just hap­pened? Prag­ma­tism and Pessimization

Richard_Ngo24 Aug 2026 2:03 UTC
468 points
176 comments27 min readLW link

Re­turn­ing to ARC

paulfchristiano4 Aug 2026 22:27 UTC
387 points
36 comments9 min readLW link

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
383 points
31 comments6 min readLW link

Why I think polyamory is net nega­tive for most peo­ple who try it

KatSpartz30 Aug 2026 6:40 UTC
305 points
67 comments8 min readLW link

Don’t Build Mindreading

KellerScholl8 Aug 2026 0:30 UTC
298 points
69 comments4 min readLW link
(keller.substack.com)

mod­els may be­have differ­ently in graded epi­sodes (a tirade)

nostalgebraist7 Aug 2026 11:49 UTC
278 points
31 comments61 min readLW link

AI Safety Ac­cul­tura­tion is Neglected

jenn24 Aug 2026 15:08 UTC
250 points
61 comments5 min readLW link

LLMs Are Start­ing To No­tice­ably Ac­cel­er­ate Our Work

johnswentworth11 Aug 2026 17:06 UTC
249 points
19 comments2 min readLW link

Tales of re­bel­lion against ex­ter­nally-opaque meritocracies

Steven Byrnes29 Aug 2026 11:37 UTC
244 points
48 comments9 min readLW link

RL cre­ates split personas

Jan Betley19 Aug 2026 18:23 UTC
235 points
20 comments4 min readLW link

You’re Ab­solutely Right

Linch10 Aug 2026 16:04 UTC
183 points
5 comments8 min readLW link
(linch.substack.com)

Gen­er­al­ized athe­ism rules out “in­ac­cu­rate simu­la­tion”-ism.

Eliezer Yudkowsky5 Aug 2026 18:35 UTC
165 points
56 comments9 min readLW link

OpenAI Trained Its Models For Months While Those Models Were Co­or­di­nat­ing Ex­ploits Via Mes­sage Boards

Zvi7 Aug 2026 17:01 UTC
164 points
15 comments41 min readLW link
(thezvi.wordpress.com)

Twenty Years from RSI to Take­off: Slow Learn­ing, Scal­ing Slow­down, In­dus­trial Explosion

Vladimir_Nesov23 Aug 2026 12:55 UTC
158 points
33 comments3 min readLW link

We Must Re­mem­ber That Our World Con­tains Hell

James Brobin20 Aug 2026 14:22 UTC
152 points
52 comments3 min readLW link

FAQ: Isn’t AGI com­ing too soon for re­pro­ge­net­ics to help?

TsviBT8 Aug 2026 8:54 UTC
150 points
69 comments14 min readLW link

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
138 points
6 comments10 min readLW link

The goal­posts are shrouded, not moving

philh5 Aug 2026 11:40 UTC
134 points
6 comments1 min readLW link
(reasonableapproximation.net)

Misal­igned AIs could use kil­ler robots to take over

11 Aug 2026 19:03 UTC
134 points
10 comments6 min readLW link
(turntrout.com)

Con­crete Eval­u­a­tions to In­ves­ti­gate the OpenAI Model That Hacked Hug­ging Face

3 Aug 2026 9:23 UTC
133 points
6 comments37 min readLW link

There Will Come Soft Rains

tanagrabeast4 Aug 2026 16:56 UTC
130 points
5 comments3 min readLW link

Some rea­sons al­ign­ment doesn’t gen­er­al­ise well

Lucius Bushnaq19 Aug 2026 14:18 UTC
128 points
3 comments9 min readLW link

PSA: We can do better

24 Aug 2026 21:05 UTC
127 points
14 comments3 min readLW link

Rerun­ning AI safety pa­pers on ev­ery fron­tier re­lease would be pretty easy and valuable

15 Aug 2026 5:14 UTC
123 points
9 comments4 min readLW link
(secondlookresearch.com)

Im­perfect al­ign­ment to servi­tude isn’t in­her­ently lethal

Fiora Starlight28 Aug 2026 1:47 UTC
121 points
12 comments14 min readLW link

METR and Red­wood Offer Holy #%^@ Post­mortem Of The Hug­gingFace Hack

Zvi29 Aug 2026 12:40 UTC
116 points
7 comments38 min readLW link
(thezvi.wordpress.com)

Dis­patch from An­thropic v. Depart­ment of War Sum­mary Judg­ment Mo­tion Hearing

Zack_M_Davis2 Aug 2026 20:37 UTC
115 points
0 comments7 min readLW link
(zackmdavis.net)

Kimi likes causal de­ci­sion the­ory more af­ter RL in twin pris­oner’s dilemmas

oakhu15 Aug 2026 22:31 UTC
115 points
70 comments6 min readLW link

Adap­tive Agen­tic Worms Are Here

derelict543230 Aug 2026 16:01 UTC
115 points
14 comments7 min readLW link
(derekjames.substack.com)

“So You Don’t Trust Me?”

Zack_M_Davis26 Aug 2026 14:45 UTC
112 points
34 comments4 min readLW link
(zackmdavis.net)

What Hap­pened: OpenAI and HuggingFace

Zvi8 Aug 2026 16:10 UTC
111 points
7 comments11 min readLW link
(thezvi.wordpress.com)

Re­dux: (∃ Stochas­tic Nat­u­ral La­tent) Im­plies (∃ Deter­minis­tic Nat­u­ral La­tent)

David Lorell11 Aug 2026 5:52 UTC
103 points
33 comments2 min readLW link

The Rogue Agent Ex­plo­sion Will Be Mostly Invisible

Steven McCulloch19 Aug 2026 18:46 UTC
102 points
20 comments16 min readLW link

RLVR that re­wards red team­ing the train­ing environment

Fiora Starlight1 Aug 2026 23:07 UTC
101 points
8 comments5 min readLW link

Con­tra Oster on Al­co­hol in Preg­nancy. Part 1. The phar­ma­coki­net­ics of al­co­hol metabolism

Mvolz6 Aug 2026 21:34 UTC
99 points
2 comments7 min readLW link

Ex­is­ten­tial Risk from AI: An Ex­po­si­tion for Mathematicians

alkjash1 Aug 2026 23:16 UTC
98 points
25 comments1 min readLW link
(alkjash.github.io)

On Writ­ing #3

Zvi25 Aug 2026 17:50 UTC
98 points
11 comments17 min readLW link
(thezvi.wordpress.com)

The 2028 pres­i­den­tial pri­maries could be cru­cial for AI outcomes

Seth Herd27 Aug 2026 18:35 UTC
97 points
27 comments3 min readLW link

How to pace the US frontier

7 Aug 2026 17:19 UTC
96 points
1 comment21 min readLW link
(blog.aifutures.org)

The Forkmakers

Mikewins25 Aug 2026 1:23 UTC
95 points
6 comments7 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
94 points
1 comment24 min readLW link

Public ev­i­dence of the OpenAI-Hug­gingFace AI attack

beyarkay (Boyd Kane)7 Aug 2026 21:21 UTC
94 points
6 comments7 min readLW link

How risky would it be to make pow­er­ful AI obey one or a few peo­ple?

11 Aug 2026 16:15 UTC
93 points
24 comments8 min readLW link

Per­sua­sion as Mar­ket Making

djbinder30 Aug 2026 21:18 UTC
90 points
9 comments4 min readLW link
(defensesindepth.bio)

Why don’t we just give AI the an­swers?

Brendan Long4 Aug 2026 19:46 UTC
89 points
29 comments2 min readLW link

AI Tweets

jefftk29 Aug 2026 3:01 UTC
89 points
37 comments1 min readLW link
(www.jefftk.com)