RSS

Some prob­lems in de­ci­sion the­ory in­cor­rectly pre­con­di­tion on policy.

Canaletto13 Aug 2026 12:04 UTC
12 points
0 comments3 min readLW link

Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems (An­thropic, Fron­tier Red Team)

Julian Bradshaw13 Aug 2026 4:10 UTC
35 points
0 comments1 min readLW link
(www.anthropic.com)

The Descen­ders and the Ab­sorbers: Two Per­spec­tives on Deep Learning

larry-dial13 Aug 2026 4:02 UTC
11 points
0 comments3 min readLW link

Free will is like temperature

Optimization Process12 Aug 2026 23:32 UTC
26 points
3 comments1 min readLW link

Every re­ward-hacked policy I tested trig­gered the OPE alarm — and so did my best hon­est one

JulesRoussel0112 Aug 2026 22:51 UTC
−5 points
0 comments48 min readLW link

Un­block­ing AI’s Con­tinual Learn­ing: Hints From How Hu­mans Learn

nimakeivan12 Aug 2026 18:28 UTC
8 points
4 comments10 min readLW link

Find­ing the Seams of Perception

jimmy12 Aug 2026 17:33 UTC
14 points
0 comments2 min readLW link
(beneathpsychology.com)

In­tro­duc­ing the Con­cep­tual Rea­son­ing Index

12 Aug 2026 17:08 UTC
71 points
7 comments5 min readLW link
(alignment.anthropic.com)

One at­ten­tion head car­ries knight forks in a chess trans­former, and here’s a new toolkit that found it.

dl2712 Aug 2026 16:48 UTC
8 points
0 comments1 min readLW link
(github.com)

The Psy­chol­ogy of Cope: Ra­tion­al­ity Is Not Re­v­ersed Irrationality

Chris_Leong12 Aug 2026 12:20 UTC
21 points
7 comments2 min readLW link

De­mon Safety

Ben Pace12 Aug 2026 11:55 UTC
46 points
4 comments1 min readLW link

The Clo­sure of the In­ter­net (Re­search Linkpost)

Dean Valentine (lc)12 Aug 2026 8:22 UTC
22 points
4 comments1 min readLW link
(arctotherium.substack.com)

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
111 points
3 comments10 min readLW link

Did the al­ign­ment com­mu­nity un­der­es­ti­mate its power?

StanislavKrym12 Aug 2026 2:56 UTC
18 points
0 comments10 min readLW link

The Age of Pluribus: One Con­sul­tant for Everyone

Dorothy Gale12 Aug 2026 1:39 UTC
9 points
0 comments4 min readLW link

We should con­sider how long mon­i­tor­ing is re­li­able for dur­ing RL

lachlan on a boat12 Aug 2026 1:22 UTC
10 points
0 comments4 min readLW link

When (and when not) LLMs can ver­bal­ize aware­ness of J-Space con­cept in­jec­tions—Ini­tial results

Ethan Garcia12 Aug 2026 1:21 UTC
8 points
0 comments15 min readLW link
(e-m-garcia.github.io)

Pa­tient Zero

LoopGameScrollMonkey12 Aug 2026 1:10 UTC
13 points
0 comments12 min readLW link

An any­time al­gorithm for mix­ing the com­putable measures

Cole Wyeth12 Aug 2026 0:57 UTC
24 points
0 comments4 min readLW link

Misal­igned AIs could use kil­ler robots to take over

11 Aug 2026 19:03 UTC
120 points
7 comments5 min readLW link