You Can’t Iter­ate to Trust­wor­thy AI Code Without Understanding

ronbodkin17 Aug 2026 23:00 UTC
8 points
0 comments10 min readLW link

Do your read­ers get an­noyed if you post ev­ery day for 30 days?

Steff17 Aug 2026 21:59 UTC
13 points
5 comments1 min readLW link

What 27 AI Safety Gen­er­al­ists Are Build­ing This Summer

17 Aug 2026 21:56 UTC
18 points
0 comments9 min readLW link

What gives you away: how LLMs form opinions of you

Cat McGee17 Aug 2026 20:59 UTC
42 points
8 comments6 min readLW link

Eval­u­at­ing Chain-of-Thought Mon­i­tora­bil­ity is Still an Open Prob­lem: Com­ments on OpenAI’s Mon­i­tora­bil­ity Evals

17 Aug 2026 19:45 UTC
20 points
3 comments18 min readLW link

Weird Re-To­k­eniza­tion, Sym­me­tries and Com­pres­sion: Re­search Agenda

17 Aug 2026 19:43 UTC
21 points
15 comments16 min readLW link

Fork Around and Find Out Part 3: In­ter­pret­ing the knight auditor

David Litman17 Aug 2026 18:11 UTC
18 points
0 comments5 min readLW link

Con­nect to your fu­ture selves

PatrickDFarley17 Aug 2026 14:57 UTC
14 points
0 comments11 min readLW link

Lat­eral Work­shop Ap­pli­ca­tions: A Case Study

17 Aug 2026 10:38 UTC
10 points
0 comments3 min readLW link
(forum.effectivealtruism.org)

Value Align­ment Is a Pseudo Con­cept. A trans­la­tion; hu­man­ity is not a sin­gle sub­ject, and al­ign­ment is not one-way.

Davidmanheim17 Aug 2026 7:44 UTC
36 points
1 comment10 min readLW link

Un­tie Squared ReLU variant

Amy_17 Aug 2026 6:15 UTC
15 points
5 comments7 min readLW link

Study Up­date: Does post-train­ing quan­ti­za­tion change welfare-rele­vant in­di­ca­tors in open-weight lan­guage mod­els?

ashesfall16 Aug 2026 21:11 UTC
8 points
0 comments7 min readLW link

Q2.5 2026 Timelines Up­date: Uplift and Revenue

16 Aug 2026 19:00 UTC
66 points
20 comments11 min readLW link
(blog.aifutures.org)

Case for Fund­ing AI Safety in Japan

mmKALLL16 Aug 2026 18:41 UTC
33 points
1 comment9 min readLW link

Will There Be an AI Hege­mon? A Men­tal Model for AI Power Concentration

simeon_c16 Aug 2026 18:41 UTC
28 points
7 comments3 min readLW link
(simeoncampos.substack.com)

The Dooms­day Ar­gu­ment is Rea­son­able and Mostly Points to Longevity

Josh Snider16 Aug 2026 17:14 UTC
29 points
12 comments6 min readLW link

Three thoughts on civil­i­sa­tional handoff

Cleo Nardo16 Aug 2026 17:12 UTC
62 points
2 comments2 min readLW link
(clattubato.substack.com)

Should Less Wrong add sub­ti­tles?

Chris_Leong16 Aug 2026 9:05 UTC
32 points
10 comments1 min readLW link

Are ques­tions al­lowed on LessWrong?

yatharth16 Aug 2026 7:04 UTC
13 points
10 comments2 min readLW link

Does Diffu­sionGemma do la­tent rea­son­ing?

16 Aug 2026 4:22 UTC
40 points
0 comments9 min readLW link

Win­ners of the Fun­da­men­tal Uncer­tainty Es­say Contest

Gordon Seidoh Worley16 Aug 2026 1:20 UTC
27 points
2 comments1 min readLW link
(www.uncertainupdates.com)

Kimi likes causal de­ci­sion the­ory more af­ter RL in twin pris­oner’s dilemmas

oakhu15 Aug 2026 22:31 UTC
115 points
70 comments6 min readLW link

Luck is a func­tion of sur­face area.

sid.the.manne@gmail.com15 Aug 2026 20:44 UTC
11 points
0 comments1 min readLW link

What if Pa­ram­e­ter Up­dates were Text?

DaemonicSigil15 Aug 2026 20:06 UTC
27 points
0 comments11 min readLW link

Us­ing Chun­ked Mon­i­tor­ing to De­tect De­cep­tion in Long Transcripts

sfereido15 Aug 2026 20:04 UTC
11 points
0 comments4 min readLW link

How To Catch a Distil­led Model

15 Aug 2026 19:09 UTC
20 points
8 comments9 min readLW link

Mom’s Ad­vice For Host­ing A Class Reunion

jenn15 Aug 2026 15:52 UTC
53 points
2 comments2 min readLW link

Learn­ing new facts can change LLM behaviour

Richard Juggins15 Aug 2026 13:46 UTC
33 points
2 comments14 min readLW link
(www.workingthroughai.com)

On Dwarkesh Pa­tel’s Pod­cast With Ryan Greenblatt

Zvi15 Aug 2026 13:00 UTC
34 points
1 comment29 min readLW link
(thezvi.wordpress.com)

Does dou­bling a user pro­file change the effect of an in­struc­tion about us­ing saved mem­o­ries?

theodorepjs15 Aug 2026 11:59 UTC
17 points
1 comment3 min readLW link

I’m start­ing a in­ter­view se­ries of peo­ple work­ing in Lean /​ for­mal meth­ods /​ math for­mal­iza­tion

Adi Baradwaj15 Aug 2026 9:15 UTC
10 points
0 comments1 min readLW link

Nu­clear physics of Alex Zhao’s com­ment for “Pac­ing the Fron­tier”

Dante Dam15 Aug 2026 9:03 UTC
17 points
0 comments11 min readLW link

Rerun­ning AI safety pa­pers on ev­ery fron­tier re­lease would be pretty easy and valuable

15 Aug 2026 5:14 UTC
123 points
9 comments4 min readLW link
(secondlookresearch.com)

Me­taphilos­o­phy II: Em­piri­cal Flywheels

interstice15 Aug 2026 2:08 UTC
18 points
12 comments5 min readLW link
(thermontology.com)

Red vs Blue, but for Evals

Morgan S15 Aug 2026 1:11 UTC
20 points
1 comment16 min readLW link

Toy Model of Ac­ti­va­tion Obfuscation

Jesse Li14 Aug 2026 23:25 UTC
16 points
0 comments7 min readLW link
(jesseli2002.github.io)

Your Agents Are Not Time Aware

14 Aug 2026 23:17 UTC
40 points
4 comments11 min readLW link

An­nounc­ing: Iliad’s New 2026 Fellowships

14 Aug 2026 22:41 UTC
37 points
0 comments1 min readLW link

Train­ing a Con­cep­tual Rea­son­ing Judge

14 Aug 2026 20:33 UTC
17 points
1 comment6 min readLW link

Why I’m Skep­ti­cal of Longtermism

James Brobin14 Aug 2026 20:17 UTC
1 point
13 comments3 min readLW link

Empty Cham­bers, Miss­ing Stairs

finitude14 Aug 2026 19:37 UTC
8 points
0 comments5 min readLW link

Do It Like Darwin

derelict543214 Aug 2026 19:11 UTC
14 points
5 comments7 min readLW link

The Day Hu­man­ity Died (Par­ody of Amer­i­can Pie)

Bentham's Bulldog14 Aug 2026 18:27 UTC
8 points
1 comment3 min readLW link

Scry­ing, Model­ing, and Nerdsnipe

Cole Wyeth14 Aug 2026 17:59 UTC
76 points
2 comments5 min readLW link

Who we’d like as regrantors

14 Aug 2026 17:40 UTC
8 points
0 comments4 min readLW link
(manifund.substack.com)

How the Amer­i­can Ex­ec­u­tive Could Con­trol AI Companies

14 Aug 2026 15:43 UTC
73 points
0 comments13 min readLW link

Open Prob­lems in Mechanis­tic In­tepretabil­ity of Biolog­i­cal AIs

Ihor Kendiukhov14 Aug 2026 15:35 UTC
21 points
0 comments2 min readLW link

V&V takes on “Pac­ing the fron­tier”

Yoav Hollander14 Aug 2026 13:25 UTC
4 points
2 comments9 min readLW link
(blog.foretellix.com)

Avoid­ing the cor­po­rate treach­er­ous turn: crowd-sourc­ing de­sign ideas

Stuart_Armstrong14 Aug 2026 13:10 UTC
26 points
14 comments1 min readLW link

Fron­tier agents don’t com­ply with stan­dards, even when in­structed to

14 Aug 2026 12:08 UTC
42 points
5 comments7 min readLW link