Free will is like temperature

Optimization Process12 Aug 2026 23:32 UTC
47 points
10 comments1 min readLW link

Im­pact mar­kets made concrete

Carol N12 Aug 2026 22:45 UTC
6 points
0 comments6 min readLW link
(manifund.substack.com)

Un­block­ing AI’s Con­tinual Learn­ing: Hints From How Hu­mans Learn

nimakeivan12 Aug 2026 18:28 UTC
3 points
23 comments10 min readLW link

Monthly Roundup #45: Au­gust 2026

Zvi12 Aug 2026 17:50 UTC
24 points
1 comment20 min readLW link
(thezvi.wordpress.com)

Find­ing the Seams of Perception

jimmy12 Aug 2026 17:33 UTC
30 points
0 comments2 min readLW link
(beneathpsychology.com)

In­tro­duc­ing the Con­cep­tual Rea­son­ing Index

12 Aug 2026 17:08 UTC
83 points
16 comments5 min readLW link
(alignment.anthropic.com)

One at­ten­tion head car­ries knight forks in a chess trans­former, and here’s a new toolkit that found it.

David Litman12 Aug 2026 16:48 UTC
8 points
0 comments1 min readLW link
(github.com)

The Psy­chol­ogy of Cope: Ra­tion­al­ity Is Not Re­v­ersed Irrationality

Chris_Leong12 Aug 2026 12:20 UTC
27 points
12 comments3 min readLW link

De­mon Safety

Ben Pace12 Aug 2026 11:55 UTC
53 points
5 comments1 min readLW link

The Clo­sure of the In­ter­net (Re­search Linkpost)

Dean Valentine12 Aug 2026 8:22 UTC
34 points
4 comments1 min readLW link
(arctotherium.substack.com)

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
138 points
6 comments10 min readLW link

Did the al­ign­ment com­mu­nity un­der­es­ti­mate its power?

StanislavKrym12 Aug 2026 2:56 UTC
23 points
0 comments10 min readLW link

The Age of Pluribus: One Con­sul­tant for Everyone

Dorothy Gale12 Aug 2026 1:39 UTC
9 points
0 comments4 min readLW link

We should con­sider how long mon­i­tor­ing is re­li­able for dur­ing RL

lachlan on a boat12 Aug 2026 1:22 UTC
20 points
1 comment4 min readLW link

When (and when not) LLMs can ver­bal­ize aware­ness of J-Space con­cept in­jec­tions—Ini­tial results

Ethan Garcia12 Aug 2026 1:21 UTC
8 points
0 comments15 min readLW link
(e-m-garcia.github.io)

Ar­gu­ments for and against (me) drop­ping out

hersheys12 Aug 2026 1:12 UTC
30 points
3 comments8 min readLW link

Pa­tient Zero

LoopGameScrollMonkey12 Aug 2026 1:10 UTC
17 points
0 comments12 min readLW link

An any­time al­gorithm for mix­ing the com­putable measures

Cole Wyeth12 Aug 2026 0:57 UTC
26 points
0 comments4 min readLW link

Var­i­ous Reflec­tions About What Hap­pened With OpenAI’s In­ter­nal Models

Zvi11 Aug 2026 22:20 UTC
63 points
1 comment30 min readLW link
(thezvi.wordpress.com)

Claude Opus 5 Just Beat My Text-Based Ad­ven­ture Game Benchmark

derelict543211 Aug 2026 20:19 UTC
22 points
4 comments5 min readLW link

[We­bi­nar] Why AI Safety is a Cap­i­tal Allo­ca­tion Prob­lem

Schizoid Rentoid11 Aug 2026 19:27 UTC
2 points
0 comments1 min readLW link

Misal­igned AIs could use kil­ler robots to take over

11 Aug 2026 19:03 UTC
134 points
10 comments6 min readLW link
(turntrout.com)

Mea­sur­ing Spu­ri­ous Cor­re­la­tions with Fea­ture Strength

egan11 Aug 2026 18:05 UTC
36 points
1 comment10 min readLW link

AI gov­er­nance work needs much bet­ter monitoring

jackultraphil11 Aug 2026 17:42 UTC
10 points
0 comments8 min readLW link
(fundinganthropalypse.com)

LLMs Are Start­ing To No­tice­ably Ac­cel­er­ate Our Work

johnswentworth11 Aug 2026 17:06 UTC
249 points
19 comments2 min readLW link

Soft­ware Is Hard

cylonator11 Aug 2026 16:26 UTC
5 points
0 comments1 min readLW link

How risky would it be to make pow­er­ful AI obey one or a few peo­ple?

11 Aug 2026 16:15 UTC
93 points
24 comments8 min readLW link

Ex­treme con­cen­tra­tion of power over ASI has non-ob­vi­ous advantages

Seth Herd11 Aug 2026 16:13 UTC
43 points
4 comments19 min readLW link

Those Who Make History

Raelifin11 Aug 2026 13:59 UTC
83 points
1 comment8 min readLW link
(open.substack.com)

See­ing things through in the age of AI

alkjash11 Aug 2026 13:14 UTC
30 points
3 comments1 min readLW link

Pro­duc­tive Sig­nal­ing: Com­pet­i­tive Soft­ware Devel­op­ment, Not Com­pet­i­tive Programming

Adam Chlipala11 Aug 2026 12:16 UTC
9 points
0 comments9 min readLW link

The Next Ecology

Eigenbraid11 Aug 2026 8:29 UTC
10 points
4 comments5 min readLW link

Re­dux: (∃ Stochas­tic Nat­u­ral La­tent) Im­plies (∃ Deter­minis­tic Nat­u­ral La­tent)

David Lorell11 Aug 2026 5:52 UTC
103 points
33 comments2 min readLW link

On us­ing crises to shift poli­ti­cal will for AI

clickyquack11 Aug 2026 4:54 UTC
17 points
0 comments4 min readLW link

Re­vived Lightweight Tran­sit Pre­dic­tions Page

jefftk11 Aug 2026 2:31 UTC
11 points
0 comments1 min readLW link
(www.jefftk.com)

Prob­ing Knowl­edge Re­cov­ery in Un­learned Models

mehnoor11 Aug 2026 2:21 UTC
7 points
0 comments7 min readLW link

Be­fore We Defer Re­search to AI: Mea­sur­ing Ap­par­ent-Suc­cess-Seeking

Keira Leal11 Aug 2026 2:21 UTC
17 points
7 comments5 min readLW link

Cana­dian Nu­clear Eman­ci­pa­tion: Canada’s role amidst global dreams of en­ergy se­cu­rity and nu­clear de­vel­op­ment

tandemocracy11 Aug 2026 2:20 UTC
8 points
0 comments3 min readLW link
(worldwithoutwater.substack.com)

A Topic De­tec­tor, Not a Lie De­tec­tor: what J-space mon­i­tor­ing ac­tu­ally tracks

Melchior de Polignac11 Aug 2026 2:18 UTC
10 points
0 comments4 min readLW link

Models in­herit the writer, not who the writer was imi­tat­ing

0Chris5R11 Aug 2026 2:16 UTC
7 points
0 comments10 min readLW link

What Claude Saw Below

Luke Nicholls11 Aug 2026 2:14 UTC
60 points
1 comment27 min readLW link

Creative math re­search by AI as the lat­est sign of the end

Mitchell_Porter11 Aug 2026 1:37 UTC
57 points
10 comments2 min readLW link

A study on in­sta­bil­ity of LLM re­sponses as a be­hav­ioral sig­na­ture of self-Refer­en­tial re­ports.

PARAS BALANI11 Aug 2026 1:01 UTC
7 points
0 comments1 min readLW link

On­line Se­quences Book Club: Begin­ners Wel­come!

gluesniffer198410 Aug 2026 22:10 UTC
1 point
0 comments1 min readLW link

The Pac­ing of the Frontier

Zvi10 Aug 2026 21:50 UTC
28 points
0 comments19 min readLW link
(thezvi.wordpress.com)

Q: Is dual-use an al­ign­ment-com­plete prob­lem?

kapedalex10 Aug 2026 21:40 UTC
10 points
1 comment1 min readLW link

Does post-train­ing quan­ti­za­tion change welfare-rele­vant in­di­ca­tors in open-weight lan­guage mod­els?

ashesfall10 Aug 2026 21:09 UTC
9 points
0 comments8 min readLW link

Claude sum­ma­rizes be­hav­ior as sig­nifi­cantly less mis­al­igned when the ac­tor is Claude vs an­other model

Ezra Newman10 Aug 2026 17:20 UTC
58 points
8 comments1 min readLW link

Off-policy hon­esty train­ing gen­er­al­izes bet­ter than on-policy hon­esty training

10 Aug 2026 16:31 UTC
20 points
0 comments9 min readLW link

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
383 points
31 comments6 min readLW link