Kimi likes causal de­ci­sion the­ory more af­ter RL in twin pris­oner’s dilemmas

oakhu15 Aug 2026 22:31 UTC
115 points
70 comments6 min readLW link

Luck is a func­tion of sur­face area.

sid.the.manne@gmail.com15 Aug 2026 20:44 UTC
11 points
0 comments1 min readLW link

What if Pa­ram­e­ter Up­dates were Text?

DaemonicSigil15 Aug 2026 20:06 UTC
27 points
0 comments11 min readLW link

Us­ing Chun­ked Mon­i­tor­ing to De­tect De­cep­tion in Long Transcripts

sfereido15 Aug 2026 20:04 UTC
11 points
0 comments4 min readLW link

How To Catch a Distil­led Model

15 Aug 2026 19:09 UTC
20 points
8 comments9 min readLW link

Mom’s Ad­vice For Host­ing A Class Reunion

jenn15 Aug 2026 15:52 UTC
53 points
2 comments2 min readLW link

Learn­ing new facts can change LLM behaviour

Richard Juggins15 Aug 2026 13:46 UTC
33 points
2 comments14 min readLW link
(www.workingthroughai.com)

On Dwarkesh Pa­tel’s Pod­cast With Ryan Greenblatt

Zvi15 Aug 2026 13:00 UTC
34 points
1 comment29 min readLW link
(thezvi.wordpress.com)

Does dou­bling a user pro­file change the effect of an in­struc­tion about us­ing saved mem­o­ries?

theodorepjs15 Aug 2026 11:59 UTC
17 points
1 comment3 min readLW link

I’m start­ing a in­ter­view se­ries of peo­ple work­ing in Lean /​ for­mal meth­ods /​ math for­mal­iza­tion

Adi Baradwaj15 Aug 2026 9:15 UTC
10 points
0 comments1 min readLW link

Nu­clear physics of Alex Zhao’s com­ment for “Pac­ing the Fron­tier”

Dante Dam15 Aug 2026 9:03 UTC
17 points
0 comments11 min readLW link

Rerun­ning AI safety pa­pers on ev­ery fron­tier re­lease would be pretty easy and valuable

15 Aug 2026 5:14 UTC
123 points
9 comments4 min readLW link
(secondlookresearch.com)

Me­taphilos­o­phy II: Em­piri­cal Flywheels

interstice15 Aug 2026 2:08 UTC
18 points
12 comments5 min readLW link
(thermontology.com)

Red vs Blue, but for Evals

Morgan S15 Aug 2026 1:11 UTC
20 points
1 comment16 min readLW link

Toy Model of Ac­ti­va­tion Obfuscation

Jesse Li14 Aug 2026 23:25 UTC
16 points
0 comments7 min readLW link
(jesseli2002.github.io)

Your Agents Are Not Time Aware

14 Aug 2026 23:17 UTC
40 points
4 comments11 min readLW link

An­nounc­ing: Iliad’s New 2026 Fellowships

14 Aug 2026 22:41 UTC
37 points
0 comments1 min readLW link

Train­ing a Con­cep­tual Rea­son­ing Judge

14 Aug 2026 20:33 UTC
17 points
1 comment6 min readLW link

Why I’m Skep­ti­cal of Longtermism

James Brobin14 Aug 2026 20:17 UTC
1 point
13 comments3 min readLW link

Empty Cham­bers, Miss­ing Stairs

finitude14 Aug 2026 19:37 UTC
8 points
0 comments5 min readLW link

Do It Like Darwin

derelict543214 Aug 2026 19:11 UTC
14 points
5 comments7 min readLW link

The Day Hu­man­ity Died (Par­ody of Amer­i­can Pie)

Bentham's Bulldog14 Aug 2026 18:27 UTC
8 points
1 comment3 min readLW link

Scry­ing, Model­ing, and Nerdsnipe

Cole Wyeth14 Aug 2026 17:59 UTC
76 points
2 comments5 min readLW link

Who we’d like as regrantors

14 Aug 2026 17:40 UTC
8 points
0 comments4 min readLW link
(manifund.substack.com)

How the Amer­i­can Ex­ec­u­tive Could Con­trol AI Companies

14 Aug 2026 15:43 UTC
73 points
0 comments13 min readLW link

Open Prob­lems in Mechanis­tic In­tepretabil­ity of Biolog­i­cal AIs

Ihor Kendiukhov14 Aug 2026 15:35 UTC
21 points
0 comments2 min readLW link

V&V takes on “Pac­ing the fron­tier”

Yoav Hollander14 Aug 2026 13:25 UTC
4 points
2 comments9 min readLW link
(blog.foretellix.com)

Avoid­ing the cor­po­rate treach­er­ous turn: crowd-sourc­ing de­sign ideas

Stuart_Armstrong14 Aug 2026 13:10 UTC
26 points
14 comments1 min readLW link

Fron­tier agents don’t com­ply with stan­dards, even when in­structed to

14 Aug 2026 12:08 UTC
42 points
5 comments7 min readLW link

What Mor­mons get right about com­mu­nity building

Jacob Brinton14 Aug 2026 7:53 UTC
36 points
2 comments4 min readLW link

Don’t for­get why learn­ing is important

Roman Ross14 Aug 2026 7:11 UTC
10 points
1 comment5 min readLW link

What If We En­forced AI Model Safety At the Level Of GPUs?

Mayowa Osibodu14 Aug 2026 6:10 UTC
1 point
5 comments4 min readLW link

Chat­ting With AIs: A Breakdown

Aditya14 Aug 2026 5:21 UTC
3 points
0 comments1 min readLW link

Fea­tures that cur­rent AIs don’t have that fu­ture AIs will have

Alexander Gietelink Oldenziel13 Aug 2026 21:49 UTC
56 points
9 comments2 min readLW link

Some Ways I Think About Eval­u­at­ing Grant Applications

sarahconstantin13 Aug 2026 21:30 UTC
61 points
0 comments10 min readLW link
(sarahconstantin.substack.com)

The Left Should Start Tak­ing AI Ca­pa­bil­ities Seriously

Alexei G13 Aug 2026 21:03 UTC
−13 points
0 comments9 min readLW link
(www.onethousandmeans.com)

Is Align­ment Even Falsifi­able? Mid­dle Align­ment, An Align­ment Tax­on­omy, and Break­ing The Prob­lem Down Into Steps

Savannah Harlan13 Aug 2026 20:57 UTC
5 points
0 comments22 min readLW link

Con­crete Gen­er­al­ist Pro­jects in AI Safety (and how to do them)

13 Aug 2026 19:26 UTC
25 points
2 comments6 min readLW link
(forum.effectivealtruism.org)

Com­par­ing Congress’s Two AI Emer­gency Shut­down Mechanisms

Philip Dowdell13 Aug 2026 18:53 UTC
13 points
0 comments8 min readLW link

Is biose­cu­rity over­sat­u­rated?

Master Chief13 Aug 2026 18:13 UTC
12 points
1 comment1 min readLW link

How to An­swer a Ques­tion Without An­swer­ing The Question

Kabir Kumar13 Aug 2026 17:06 UTC
37 points
3 comments1 min readLW link

How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
772 points
109 comments12 min readLW link

AI #181: As­tra Goes Cy­ber Critical

Zvi13 Aug 2026 15:20 UTC
39 points
1 comment52 min readLW link
(thezvi.wordpress.com)

Au­to­mated al­ign­ment runs are hard to study!

13 Aug 2026 15:04 UTC
67 points
4 comments9 min readLW link

Longter­mism Seems Like A Religion

James Brobin13 Aug 2026 14:43 UTC
2 points
8 comments3 min readLW link

Defense Against the De­cep­tive Arts

Kabir Kumar13 Aug 2026 14:21 UTC
11 points
11 comments1 min readLW link

Some prob­lems in de­ci­sion the­ory in­cor­rectly pre­con­di­tion on policy.

Canaletto13 Aug 2026 12:04 UTC
10 points
4 comments3 min readLW link

Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems (An­thropic, Fron­tier Red Team)

Julian Bradshaw13 Aug 2026 4:10 UTC
42 points
3 comments1 min readLW link
(www.anthropic.com)

The Descen­ders and the Ab­sorbers: Two Per­spec­tives on Deep Learning

larry-dial13 Aug 2026 4:02 UTC
11 points
0 comments3 min readLW link

OC ACXLW Meetup #118 — The 40% Prob­lem & The Field That Re­fused to Die

Michael Michalchik13 Aug 2026 3:25 UTC
1 point
0 comments9 min readLW link