Cui bono? ChatGPT-4o shows non-de­cep­tive strate­gic per­sua­sion: a Proof-of-Con­cept study

Pieter Barkema18 Aug 2026 23:25 UTC
7 points
0 comments5 min readLW link

Whack-a-mole with a bro­ken ham­mer: does a model in­ter­nally track its au­toma­ton state?

star2vec18 Aug 2026 23:25 UTC
7 points
0 comments8 min readLW link

GPT-2′s IOI be­hav­ior is defined where the pa­per’s al­gorithm isn’t

Zach Allen18 Aug 2026 23:23 UTC
7 points
0 comments2 min readLW link

Put­ting your mask on first

emiliob18 Aug 2026 23:23 UTC
1 point
0 comments1 min readLW link
(emiliobarkett.github.io)

Tran­shu­man­ism & Hu­man En­hance­ment Meetup (On­line)

nxzt18 Aug 2026 23:22 UTC
1 point
0 comments1 min readLW link

Goal-di­rected op­ti­mi­sa­tion re­quires in­for­ma­tion about the world

dubinsschwarz18 Aug 2026 23:22 UTC
13 points
0 comments12 min readLW link

AI Sand­bag­ging (w/​ In­spect)

Leo H18 Aug 2026 23:20 UTC
8 points
0 comments5 min readLW link

An­thropic Risk Re­port: Au­gust 2026

Zvi18 Aug 2026 20:40 UTC
48 points
3 comments42 min readLW link
(thezvi.wordpress.com)

Fund­ing For­mal Meth­ods for the Cyberpocalypse

Max von Hippel18 Aug 2026 16:45 UTC
22 points
8 comments9 min readLW link

Policy ca­reer plan­ning in the age of im­mi­nent superintelligence

Peter Wildeford18 Aug 2026 14:33 UTC
70 points
5 comments6 min readLW link

Au­to­matic Pro­gram­ming Should Be More Like SQL

Adam Chlipala18 Aug 2026 12:29 UTC
7 points
0 comments12 min readLW link

AI Se­cu­rity is Harm Reduction

Quinn18 Aug 2026 12:28 UTC
39 points
3 comments1 min readLW link

LASR Labs Win­ter 2027 ap­pli­ca­tions are open!

18 Aug 2026 10:12 UTC
37 points
1 comment4 min readLW link

Price re­cur­sion is the ra­tio­nal the­ory of reward

Abhimanyu Pallavi Sudhir18 Aug 2026 3:22 UTC
6 points
0 comments9 min readLW link

Nat­u­ral In­de­pen­dence Incentives

jefftk18 Aug 2026 2:40 UTC
55 points
0 comments2 min readLW link
(www.jefftk.com)

Misal­igned In­cen­tives in Pause Scenarios

18 Aug 2026 1:55 UTC
48 points
13 comments17 min readLW link

Inoc­u­late Every­thing: All of Pre­train­ing and RL

davids18 Aug 2026 1:28 UTC
9 points
0 comments10 min readLW link
(zenodo.org)

For Claude, ca­pa­bil­ity and dis­prefer­ring CDT are the ~same thing. Much more so than for GPT.

18 Aug 2026 1:25 UTC
50 points
13 comments1 min readLW link

You Can’t Iter­ate to Trust­wor­thy AI Code Without Understanding

ronbodkin17 Aug 2026 23:00 UTC
8 points
0 comments10 min readLW link

Do your read­ers get an­noyed if you post ev­ery day for 30 days?

Steff17 Aug 2026 21:59 UTC
13 points
5 comments1 min readLW link

What 27 AI Safety Gen­er­al­ists Are Build­ing This Summer

17 Aug 2026 21:56 UTC
18 points
0 comments9 min readLW link

What gives you away: how LLMs form opinions of you

Cat McGee17 Aug 2026 20:59 UTC
42 points
8 comments6 min readLW link

Eval­u­at­ing Chain-of-Thought Mon­i­tora­bil­ity is Still an Open Prob­lem: Com­ments on OpenAI’s Mon­i­tora­bil­ity Evals

17 Aug 2026 19:45 UTC
20 points
3 comments18 min readLW link

Weird Re-To­k­eniza­tion, Sym­me­tries and Com­pres­sion: Re­search Agenda

17 Aug 2026 19:43 UTC
21 points
15 comments16 min readLW link

Fork Around and Find Out Part 3: In­ter­pret­ing the knight auditor

David Litman17 Aug 2026 18:11 UTC
18 points
0 comments5 min readLW link

Con­nect to your fu­ture selves

PatrickDFarley17 Aug 2026 14:57 UTC
14 points
0 comments11 min readLW link

Lat­eral Work­shop Ap­pli­ca­tions: A Case Study

17 Aug 2026 10:38 UTC
7 points
0 comments3 min readLW link
(forum.effectivealtruism.org)

Value Align­ment Is a Pseudo Con­cept. A trans­la­tion; hu­man­ity is not a sin­gle sub­ject, and al­ign­ment is not one-way.

Davidmanheim17 Aug 2026 7:44 UTC
36 points
1 comment10 min readLW link

Un­tie Squared ReLU variant

Amy_17 Aug 2026 6:15 UTC
15 points
5 comments7 min readLW link

Study Up­date: Does post-train­ing quan­ti­za­tion change welfare-rele­vant in­di­ca­tors in open-weight lan­guage mod­els?

ashesfall16 Aug 2026 21:11 UTC
8 points
0 comments7 min readLW link

Q2.5 2026 Timelines Up­date: Uplift and Revenue

16 Aug 2026 19:00 UTC
66 points
20 comments11 min readLW link
(blog.aifutures.org)

Case for Fund­ing AI Safety in Japan

mmKALLL16 Aug 2026 18:41 UTC
33 points
1 comment9 min readLW link

Will There Be an AI Hege­mon? A Men­tal Model for AI Power Concentration

simeon_c16 Aug 2026 18:41 UTC
28 points
7 comments3 min readLW link
(simeoncampos.substack.com)

The Dooms­day Ar­gu­ment is Rea­son­able and Mostly Points to Longevity

Josh Snider16 Aug 2026 17:14 UTC
29 points
11 comments6 min readLW link

Three thoughts on civil­i­sa­tional handoff

Cleo Nardo16 Aug 2026 17:12 UTC
62 points
2 comments2 min readLW link
(clattubato.substack.com)

Should Less Wrong add sub­ti­tles?

Chris_Leong16 Aug 2026 9:05 UTC
32 points
10 comments1 min readLW link

Are ques­tions al­lowed on LessWrong?

yatharth16 Aug 2026 7:04 UTC
13 points
10 comments2 min readLW link

Does Diffu­sionGemma do la­tent rea­son­ing?

16 Aug 2026 4:22 UTC
40 points
0 comments9 min readLW link

Win­ners of the Fun­da­men­tal Uncer­tainty Es­say Contest

Gordon Seidoh Worley16 Aug 2026 1:20 UTC
27 points
2 comments1 min readLW link
(www.uncertainupdates.com)

Kimi likes causal de­ci­sion the­ory more af­ter RL in twin pris­oner’s dilemmas

oakhu15 Aug 2026 22:31 UTC
115 points
70 comments6 min readLW link

Luck is a func­tion of sur­face area.

sid.the.manne@gmail.com15 Aug 2026 20:44 UTC
11 points
0 comments1 min readLW link

What if Pa­ram­e­ter Up­dates were Text?

DaemonicSigil15 Aug 2026 20:06 UTC
27 points
0 comments11 min readLW link

Us­ing Chun­ked Mon­i­tor­ing to De­tect De­cep­tion in Long Transcripts

sfereido15 Aug 2026 20:04 UTC
11 points
0 comments4 min readLW link

How To Catch a Distil­led Model

15 Aug 2026 19:09 UTC
20 points
8 comments9 min readLW link

Mom’s Ad­vice For Host­ing A Class Reunion

jenn15 Aug 2026 15:52 UTC
53 points
2 comments2 min readLW link

Learn­ing new facts can change LLM behaviour

Richard Juggins15 Aug 2026 13:46 UTC
33 points
2 comments14 min readLW link
(www.workingthroughai.com)

On Dwarkesh Pa­tel’s Pod­cast With Ryan Greenblatt

Zvi15 Aug 2026 13:00 UTC
34 points
1 comment29 min readLW link
(thezvi.wordpress.com)

Does dou­bling a user pro­file change the effect of an in­struc­tion about us­ing saved mem­o­ries?

theodorepjs15 Aug 2026 11:59 UTC
17 points
1 comment3 min readLW link

I’m start­ing a in­ter­view se­ries of peo­ple work­ing in Lean /​ for­mal meth­ods /​ math for­mal­iza­tion

Adi Baradwaj15 Aug 2026 9:15 UTC
10 points
0 comments1 min readLW link

Nu­clear physics of Alex Zhao’s com­ment for “Pac­ing the Fron­tier”

Dante Dam15 Aug 2026 9:03 UTC
17 points
0 comments11 min readLW link