Var­i­ous Reflec­tions About What Hap­pened With OpenAI’s In­ter­nal Models

Zvi11 Aug 2026 22:20 UTC
63 points
1 comment30 min readLW link
(thezvi.wordpress.com)

Claude Opus 5 Just Beat My Text-Based Ad­ven­ture Game Benchmark

derelict543211 Aug 2026 20:19 UTC
22 points
4 comments5 min readLW link

[We­bi­nar] Why AI Safety is a Cap­i­tal Allo­ca­tion Prob­lem

Schizoid Rentoid11 Aug 2026 19:27 UTC
2 points
0 comments1 min readLW link

Misal­igned AIs could use kil­ler robots to take over

11 Aug 2026 19:03 UTC
134 points
10 comments6 min readLW link
(turntrout.com)

Mea­sur­ing Spu­ri­ous Cor­re­la­tions with Fea­ture Strength

egan11 Aug 2026 18:05 UTC
36 points
1 comment10 min readLW link

AI gov­er­nance work needs much bet­ter monitoring

jackultraphil11 Aug 2026 17:42 UTC
10 points
0 comments8 min readLW link
(fundinganthropalypse.com)

LLMs Are Start­ing To No­tice­ably Ac­cel­er­ate Our Work

johnswentworth11 Aug 2026 17:06 UTC
249 points
19 comments2 min readLW link

Soft­ware Is Hard

cylonator11 Aug 2026 16:26 UTC
5 points
0 comments1 min readLW link

How risky would it be to make pow­er­ful AI obey one or a few peo­ple?

11 Aug 2026 16:15 UTC
93 points
24 comments8 min readLW link

Ex­treme con­cen­tra­tion of power over ASI has non-ob­vi­ous advantages

Seth Herd11 Aug 2026 16:13 UTC
43 points
4 comments19 min readLW link

Those Who Make History

Raelifin11 Aug 2026 13:59 UTC
83 points
1 comment8 min readLW link
(open.substack.com)

See­ing things through in the age of AI

alkjash11 Aug 2026 13:14 UTC
30 points
3 comments1 min readLW link

Pro­duc­tive Sig­nal­ing: Com­pet­i­tive Soft­ware Devel­op­ment, Not Com­pet­i­tive Programming

Adam Chlipala11 Aug 2026 12:16 UTC
9 points
0 comments9 min readLW link

The Next Ecology

Eigenbraid11 Aug 2026 8:29 UTC
10 points
4 comments5 min readLW link

Re­dux: (∃ Stochas­tic Nat­u­ral La­tent) Im­plies (∃ Deter­minis­tic Nat­u­ral La­tent)

David Lorell11 Aug 2026 5:52 UTC
103 points
33 comments2 min readLW link

On us­ing crises to shift poli­ti­cal will for AI

clickyquack11 Aug 2026 4:54 UTC
17 points
0 comments4 min readLW link

Re­vived Lightweight Tran­sit Pre­dic­tions Page

jefftk11 Aug 2026 2:31 UTC
11 points
0 comments1 min readLW link
(www.jefftk.com)

Prob­ing Knowl­edge Re­cov­ery in Un­learned Models

mehnoor11 Aug 2026 2:21 UTC
7 points
0 comments7 min readLW link

Be­fore We Defer Re­search to AI: Mea­sur­ing Ap­par­ent-Suc­cess-Seeking

Keira Leal11 Aug 2026 2:21 UTC
17 points
7 comments5 min readLW link

Cana­dian Nu­clear Eman­ci­pa­tion: Canada’s role amidst global dreams of en­ergy se­cu­rity and nu­clear de­vel­op­ment

tandemocracy11 Aug 2026 2:20 UTC
8 points
0 comments3 min readLW link
(worldwithoutwater.substack.com)

A Topic De­tec­tor, Not a Lie De­tec­tor: what J-space mon­i­tor­ing ac­tu­ally tracks

Melchior de Polignac11 Aug 2026 2:18 UTC
10 points
0 comments4 min readLW link

Models in­herit the writer, not who the writer was imi­tat­ing

0Chris5R11 Aug 2026 2:16 UTC
7 points
0 comments10 min readLW link

What Claude Saw Below

Luke Nicholls11 Aug 2026 2:14 UTC
60 points
1 comment27 min readLW link

Creative math re­search by AI as the lat­est sign of the end

Mitchell_Porter11 Aug 2026 1:37 UTC
57 points
10 comments2 min readLW link

A study on in­sta­bil­ity of LLM re­sponses as a be­hav­ioral sig­na­ture of self-Refer­en­tial re­ports.

PARAS BALANI11 Aug 2026 1:01 UTC
7 points
0 comments1 min readLW link

On­line Se­quences Book Club: Begin­ners Wel­come!

gluesniffer198410 Aug 2026 22:10 UTC
1 point
0 comments1 min readLW link

The Pac­ing of the Frontier

Zvi10 Aug 2026 21:50 UTC
28 points
0 comments19 min readLW link
(thezvi.wordpress.com)

Q: Is dual-use an al­ign­ment-com­plete prob­lem?

kapedalex10 Aug 2026 21:40 UTC
10 points
1 comment1 min readLW link

Does post-train­ing quan­ti­za­tion change welfare-rele­vant in­di­ca­tors in open-weight lan­guage mod­els?

ashesfall10 Aug 2026 21:09 UTC
9 points
0 comments8 min readLW link

Claude sum­ma­rizes be­hav­ior as sig­nifi­cantly less mis­al­igned when the ac­tor is Claude vs an­other model

Ezra Newman10 Aug 2026 17:20 UTC
58 points
8 comments1 min readLW link

Off-policy hon­esty train­ing gen­er­al­izes bet­ter than on-policy hon­esty training

10 Aug 2026 16:31 UTC
20 points
0 comments9 min readLW link

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
383 points
31 comments6 min readLW link

Co­er­cion and De­cep­tion in AI-to-AI Management

10 Aug 2026 16:13 UTC
12 points
0 comments8 min readLW link
(compassionalignedml.substack.com)

You’re Ab­solutely Right

Linch10 Aug 2026 16:04 UTC
183 points
5 comments8 min readLW link
(linch.substack.com)

Book Re­view: The In­finity Machine

Mikewins10 Aug 2026 16:01 UTC
10 points
0 comments1 min readLW link

On De­moc­ra­tiz­ing ASI to Pre­serve Civil Liberties

MichaelDickens10 Aug 2026 12:08 UTC
43 points
5 comments3 min readLW link

Is Eval Gam­ing Down­stream of Ver­bal­ized Eval Aware­ness? Not when it’s re­flex­ive.

Kieron Kretschmar10 Aug 2026 8:24 UTC
14 points
2 comments11 min readLW link

Is it eth­i­cal to work on gen­eral-pur­pose robots given the risk of to­tal­i­tar­i­anism?

Master Chief10 Aug 2026 6:40 UTC
22 points
5 comments1 min readLW link

How to be an AI safety re­search engineer

Ruben Castaing10 Aug 2026 5:32 UTC
11 points
0 comments6 min readLW link

Hiring Vibe-wran­gler Match­mak­ing Thread

Raemon10 Aug 2026 3:53 UTC
26 points
1 comment3 min readLW link

How to get an­swers to ques­tions that con­fuse you (maybe)

Elijah10 Aug 2026 1:35 UTC
5 points
0 comments5 min readLW link

The Agen­tic Clusterfuck

Chapin Lenthall-Cleary10 Aug 2026 1:27 UTC
51 points
15 comments1 min readLW link

Over­think­ing: Am­plify­ing rea­son­ing weights makes mod­els re­veal their secrets

9 Aug 2026 22:06 UTC
22 points
0 comments7 min readLW link
(arxiv.org)

AI-am­plified demo­cratic back­slid­ing: an exploration

9 Aug 2026 22:05 UTC
17 points
3 comments6 min readLW link

The Apoca­lyp­tic Ar­rival of Truth

Caleb Biddulph9 Aug 2026 21:32 UTC
43 points
2 comments1 min readLW link

“Com­mu­nity Notes” re­s­olu­tion for vague predictions

Raemon9 Aug 2026 19:51 UTC
43 points
22 comments1 min readLW link

Ten Thou­sand Cy­ber Labs for Train­ing & Eval

TheVinci9 Aug 2026 17:57 UTC
10 points
3 comments3 min readLW link

A challenge: Can you make an LLM fol­low these in­struc­tions?

Steff9 Aug 2026 16:50 UTC
18 points
7 comments6 min readLW link

What just hap­pened? A ret­ro­spec­tive of AI alignment

Richard_Ngo9 Aug 2026 15:58 UTC
640 points
152 comments16 min readLW link

Who does the con­fess­ing, and will they con­fess to anything

Abhishek Mishra9 Aug 2026 13:34 UTC
7 points
0 comments8 min readLW link