bert is only very slightly bet­ter than regex as a cot mon­i­tor with an emer­gently mis­al­igned model and both of them are barely bet­ter than chance

nesiacel26 Aug 2026 22:27 UTC
16 points
0 comments2 min readLW link

The cen­tral fal­lacy: the sec­ond-worst ar­gu­ment in the world

Connor Williams26 Aug 2026 21:00 UTC
12 points
3 comments6 min readLW link
(connorsscratchpad.substack.com)

Your Agent’s Trace Prob­a­bly Can­not Tell You Who Ap­proved a Tool Call

Jaswanth Alkur26 Aug 2026 20:27 UTC
6 points
0 comments3 min readLW link

Mon­day for Future

Ondřej Lukeš26 Aug 2026 20:20 UTC
1 point
0 comments1 min readLW link

Do tab­u­lar foun­da­tion mod­els re­pair them­selves?

Han Xiao26 Aug 2026 20:19 UTC
6 points
0 comments8 min readLW link

Be­ing Neu­rotic about Fer­til­ity: Notes from the 2026 Re­pro­duc­tive Fron­tiers Conference

boba_girl26 Aug 2026 20:18 UTC
55 points
4 comments8 min readLW link

Clear­ing My Foggy Glasses

adamShimi26 Aug 2026 20:18 UTC
36 points
2 comments12 min readLW link
(formethods.substack.com)

My MATS 11.0 Ap­pli­ca­tion Experience

Evan Conway26 Aug 2026 20:17 UTC
11 points
1 comment7 min readLW link
(evanjayconway.com)

Brief in­de­pen­dent in­ves­ti­ga­tion of agents’ be­hav­ior, rea­son­ing and col­lab­o­ra­tion in the OpenAI /​ Hug­ging Face hack­ing incident

26 Aug 2026 19:40 UTC
592 points
67 comments3 min readLW link
(metr.org)

The Brain Preser­va­tion Sum­mit—Septem­ber 19-20th at Sparks Brain Preser­va­tion in Salem, OR

Andy_McKenzie26 Aug 2026 19:35 UTC
21 points
0 comments2 min readLW link
(sparksbrain.org)

Against Modesty’s Bailey

Zvi26 Aug 2026 18:20 UTC
74 points
6 comments14 min readLW link
(thezvi.wordpress.com)

Where Did D Go? A Gap Between ARC’s Mo­ti­va­tion and Its Formalism

Zach Allen26 Aug 2026 18:05 UTC
10 points
0 comments4 min readLW link
(www.lesswrong.com)

Devel­op­ing frames

Richard_Ngo26 Aug 2026 17:41 UTC
30 points
0 comments10 min readLW link

[Macroa­gents] 1. The Macroa­gent Ontology

Towards_Keeperhood26 Aug 2026 16:48 UTC
20 points
0 comments6 min readLW link

An­nounc­ing the MATS Res­i­dency: A New Path for Ex­pe­rienced Re­searchers Work­ing on AI Safety, or Mov­ing Into it

Raj Thimmiah26 Aug 2026 16:07 UTC
20 points
0 comments4 min readLW link

LLMs are Adap­ta­tion-Ex­e­cuters, not Un­der­stand­ing-Maximizers

olehif26 Aug 2026 15:48 UTC
7 points
0 comments1 min readLW link

“So You Don’t Trust Me?”

Zack_M_Davis26 Aug 2026 14:45 UTC
112 points
34 comments4 min readLW link
(zackmdavis.net)

Safety Cases We Can Check — To­gether!

Audrey Tang26 Aug 2026 13:56 UTC
13 points
0 comments1 min readLW link
(au.civic.ai)

Gem­ini 2.5 Pro in the AI Village as a Nat­u­ral Case Study of Com­pound­ing Misalignment

26 Aug 2026 8:51 UTC
13 points
0 comments11 min readLW link
(aivillageblog.substack.com)

Would this re­search out­line be in­ter­est­ing? I would like to hear your opinions.

R U26 Aug 2026 6:12 UTC
3 points
2 comments9 min readLW link

We need more em­piri­cal re­search on biolog­i­cal AI model safety

Noga Aharony26 Aug 2026 3:07 UTC
14 points
2 comments5 min readLW link

AI safety through the lens of un­cer­tainty quan­tifi­ca­tion (UQ).

arradiat26 Aug 2026 3:07 UTC
9 points
2 comments2 min readLW link

Short Time Outs

jefftk26 Aug 2026 0:01 UTC
26 points
4 comments2 min readLW link
(www.jefftk.com)

When There Are No Experts

J Bostock25 Aug 2026 22:47 UTC
51 points
7 comments4 min readLW link

In­te­grated Strate­gic Fore­cast­ing: A New RAND x Me­tac­u­lus Methodology

25 Aug 2026 20:05 UTC
22 points
1 comment2 min readLW link
(www.metaculus.com)

Sam­pura Re­search: Hu­man-AI Com­ple­men­tar­ity for Alignment

25 Aug 2026 18:45 UTC
17 points
2 comments5 min readLW link
(sampura.org)

On Writ­ing #3

Zvi25 Aug 2026 17:50 UTC
98 points
11 comments17 min readLW link
(thezvi.wordpress.com)

The Reified Refer­ent: How the Sat­u­ra­tion View Values the Taste of Some­one Who Does Not Exist

Ronen Bar25 Aug 2026 15:26 UTC
0 points
12 comments6 min readLW link
(forum.effectivealtruism.org)

AI Align­ment at Which Ab­strac­tion Level?

Adam Chlipala25 Aug 2026 12:37 UTC
5 points
1 comment9 min readLW link

The Forkmakers

Mikewins25 Aug 2026 1:23 UTC
96 points
6 comments7 min readLW link

An AI4Chemist’s per­spec­tive on biosecurity

Ana Leonescu25 Aug 2026 1:07 UTC
13 points
0 comments8 min readLW link
(substack.com)

Magaz­ine Fundraising

Elan Kluger25 Aug 2026 1:06 UTC
24 points
0 comments1 min readLW link

Rogue AI Agents: Is Sur­face-Level Mon­i­tor­ing Enough?

Hariom Tatsat25 Aug 2026 1:05 UTC
8 points
0 comments5 min readLW link

Dual-Layer Ap­proach to Miti­gat­ing AI-Gen­er­ated Biolog­i­cal Threats

greeshma ranjeev25 Aug 2026 1:02 UTC
7 points
0 comments6 min readLW link

Your Eval­u­a­tion’s Fake names Should Be Un­claimable ,Not Merely Used

Jaswanth Alkur25 Aug 2026 1:01 UTC
26 points
2 comments2 min readLW link

My AI Syl­labus Policy

Vaughn Papenhausen24 Aug 2026 21:36 UTC
26 points
2 comments4 min readLW link
(vaughnpapenhausen.substack.com)

PSA: We can do better

24 Aug 2026 21:05 UTC
127 points
14 comments3 min readLW link

Distil­la­tion of the AI2040 Align­ment Roadmap

Towards_Keeperhood24 Aug 2026 18:20 UTC
24 points
0 comments23 min readLW link

The Amer­i­can Peo­ple Really Hate Data Centers

Zvi24 Aug 2026 18:10 UTC
31 points
2 comments15 min readLW link
(thezvi.wordpress.com)

Where have organoids ac­tu­ally been use­ful?

Abhishaike Mahajan24 Aug 2026 18:01 UTC
32 points
0 comments19 min readLW link

AI Safety Ac­cul­tura­tion is Neglected

jenn24 Aug 2026 15:08 UTC
250 points
61 comments5 min readLW link

En­dors­ing Burhan Azeem for State Senate

jefftk24 Aug 2026 12:10 UTC
23 points
0 comments2 min readLW link
(www.jefftk.com)

LLMs could con­trol their host ma­chines by ex­ploit­ing in­fer­ence engines

beyarkay (Boyd Kane)24 Aug 2026 9:45 UTC
72 points
7 comments4 min readLW link
(boydkane.com)

I Can Still Hear You Sayin’

LoopGameScrollMonkey24 Aug 2026 5:21 UTC
1 point
0 comments3 min readLW link

What just hap­pened? Prag­ma­tism and Pessimization

Richard_Ngo24 Aug 2026 2:03 UTC
473 points
176 comments27 min readLW link

In search of nat­u­ral features

Dmitry Vaintrob23 Aug 2026 22:15 UTC
55 points
3 comments17 min readLW link

PSA: There’s a third op­tion in the “mea­sure prob­lem”

Elias Schmied23 Aug 2026 16:41 UTC
48 points
29 comments3 min readLW link

Utilities as Le­gen­dre du­als of probabilities

Fernando Rosas23 Aug 2026 16:26 UTC
69 points
11 comments6 min readLW link

Twenty Years from RSI to Take­off: Slow Learn­ing, Scal­ing Slow­down, In­dus­trial Explosion

Vladimir_Nesov23 Aug 2026 12:55 UTC
160 points
33 comments3 min readLW link

How to over­haul the bro­ken re­view system

RobinHa23 Aug 2026 9:14 UTC
19 points
1 comment13 min readLW link