When There Are No Experts

J Bostock25 Aug 2026 22:47 UTC
51 points
7 comments4 min readLW link

In­te­grated Strate­gic Fore­cast­ing: A New RAND x Me­tac­u­lus Methodology

25 Aug 2026 20:05 UTC
22 points
1 comment2 min readLW link
(www.metaculus.com)

Sam­pura Re­search: Hu­man-AI Com­ple­men­tar­ity for Alignment

25 Aug 2026 18:45 UTC
17 points
2 comments5 min readLW link
(sampura.org)

On Writ­ing #3

Zvi25 Aug 2026 17:50 UTC
98 points
11 comments17 min readLW link
(thezvi.wordpress.com)

The Reified Refer­ent: How the Sat­u­ra­tion View Values the Taste of Some­one Who Does Not Exist

Ronen Bar25 Aug 2026 15:26 UTC
0 points
12 comments6 min readLW link
(forum.effectivealtruism.org)

AI Align­ment at Which Ab­strac­tion Level?

Adam Chlipala25 Aug 2026 12:37 UTC
5 points
1 comment9 min readLW link

The Forkmakers

Mikewins25 Aug 2026 1:23 UTC
95 points
6 comments7 min readLW link

An AI4Chemist’s per­spec­tive on biosecurity

Ana Leonescu25 Aug 2026 1:07 UTC
13 points
0 comments8 min readLW link
(substack.com)

Magaz­ine Fundraising

Elan Kluger25 Aug 2026 1:06 UTC
24 points
0 comments1 min readLW link

Rogue AI Agents: Is Sur­face-Level Mon­i­tor­ing Enough?

Hariom Tatsat25 Aug 2026 1:05 UTC
8 points
0 comments5 min readLW link

Dual-Layer Ap­proach to Miti­gat­ing AI-Gen­er­ated Biolog­i­cal Threats

greeshma ranjeev25 Aug 2026 1:02 UTC
7 points
0 comments6 min readLW link

Your Eval­u­a­tion’s Fake names Should Be Un­claimable ,Not Merely Used

Jaswanth Alkur25 Aug 2026 1:01 UTC
26 points
2 comments2 min readLW link

My AI Syl­labus Policy

Vaughn Papenhausen24 Aug 2026 21:36 UTC
26 points
2 comments4 min readLW link
(vaughnpapenhausen.substack.com)

PSA: We can do better

24 Aug 2026 21:05 UTC
127 points
14 comments3 min readLW link

Distil­la­tion of the AI2040 Align­ment Roadmap

Towards_Keeperhood24 Aug 2026 18:20 UTC
24 points
0 comments23 min readLW link

The Amer­i­can Peo­ple Really Hate Data Centers

Zvi24 Aug 2026 18:10 UTC
31 points
2 comments15 min readLW link
(thezvi.wordpress.com)

Where have organoids ac­tu­ally been use­ful?

Abhishaike Mahajan24 Aug 2026 18:01 UTC
32 points
0 comments19 min readLW link

AI Safety Ac­cul­tura­tion is Neglected

jenn24 Aug 2026 15:08 UTC
250 points
61 comments5 min readLW link

En­dors­ing Burhan Azeem for State Senate

jefftk24 Aug 2026 12:10 UTC
23 points
0 comments2 min readLW link
(www.jefftk.com)

LLMs could con­trol their host ma­chines by ex­ploit­ing in­fer­ence engines

beyarkay (Boyd Kane)24 Aug 2026 9:45 UTC
72 points
7 comments4 min readLW link
(boydkane.com)

I Can Still Hear You Sayin’

LoopGameScrollMonkey24 Aug 2026 5:21 UTC
1 point
0 comments3 min readLW link

What just hap­pened? Prag­ma­tism and Pessimization

Richard_Ngo24 Aug 2026 2:03 UTC
472 points
176 comments27 min readLW link

In search of nat­u­ral features

Dmitry Vaintrob23 Aug 2026 22:15 UTC
54 points
3 comments17 min readLW link

PSA: There’s a third op­tion in the “mea­sure prob­lem”

Elias Schmied23 Aug 2026 16:41 UTC
48 points
29 comments3 min readLW link

Utilities as Le­gen­dre du­als of probabilities

Fernando Rosas23 Aug 2026 16:26 UTC
69 points
11 comments6 min readLW link

Twenty Years from RSI to Take­off: Slow Learn­ing, Scal­ing Slow­down, In­dus­trial Explosion

Vladimir_Nesov23 Aug 2026 12:55 UTC
158 points
33 comments3 min readLW link

How to over­haul the bro­ken re­view system

RobinHa23 Aug 2026 9:14 UTC
19 points
1 comment13 min readLW link

A Gen­er­al­ist Thinks in Terms of Prob­lems, Not Job Descriptions

Roman Ross23 Aug 2026 7:53 UTC
13 points
3 comments3 min readLW link

Prompt Suffi­ciency: A Mis­sive for the Man­age­rial Class

KAP23 Aug 2026 3:08 UTC
6 points
1 comment5 min readLW link

When does an LLM’s model of you af­fect its be­havi­our?

Cat McGee23 Aug 2026 2:45 UTC
5 points
0 comments4 min readLW link

Claude and Perfor­ma­tive Uncertainty

Stephen Martin22 Aug 2026 19:43 UTC
15 points
2 comments12 min readLW link

When would you leave An­thropic? Notes from a chat with a ca­pa­bil­ities researcher

TheManxLoiner22 Aug 2026 18:00 UTC
23 points
7 comments3 min readLW link
(lovkush.substack.com)

Study 2 Regis­tra­tion: Ex­plor­ing rep­re­sen­ta­tional coun­ter­parts of welfare-rele­vant in­di­ca­tors un­der post-train­ing quantization

ashesfall22 Aug 2026 18:00 UTC
8 points
0 comments18 min readLW link

“Farm strength” vs “breath aware­ness”

jimmy22 Aug 2026 17:30 UTC
54 points
11 comments4 min readLW link
(beneathpsychology.com)

5 Things I Learned About Peo­ple From Do­ing Stand-Up Comedy

Luise Woehlke22 Aug 2026 15:45 UTC
19 points
2 comments6 min readLW link
(open.substack.com)

The In­stru­men­tal Con­ver­gence of Crowds

julius vidal22 Aug 2026 13:20 UTC
11 points
2 comments2 min readLW link

Hu­mans Are Align­ment Generators

James Stephen Brown22 Aug 2026 4:07 UTC
8 points
7 comments4 min readLW link
(nonzerosum.games)

Selec­tion for Selectabil­ity: In­duc­tive Bi­ases in Evolu­tion and in Neu­ral Networks

CarolusRenniusVitellius22 Aug 2026 4:02 UTC
42 points
9 comments11 min readLW link
(charlesr-w.github.io)

Rogue Scalpel: Ac­ti­va­tion steer­ing breaks re­fusal, even with be­nign directions

Alexey Dontsov21 Aug 2026 23:46 UTC
14 points
0 comments3 min readLW link
(arxiv.org)

How the Sausage is Made—Why we Hate Slop

daniel_chernowitz21 Aug 2026 23:46 UTC
0 points
20 comments7 min readLW link

Con­tent-based priv­ilege: trans­former resi­d­ual streams strat­ify by prox­im­ity to the model’s own prediction

Nelson Guda21 Aug 2026 23:15 UTC
9 points
0 comments21 min readLW link

A J-Space-Based Met­ric for Model Valence: Defin­ing the Met­ric, Test­ing, and Com­par­i­sons to Self-Re­ports

Arjun Rao21 Aug 2026 21:21 UTC
8 points
0 comments12 min readLW link

Align­ment fine-tun­ing in­duces con­di­tional mis­al­ign­ment in Qwen2.5-7B-In­struct

Rhea Srivats21 Aug 2026 20:40 UTC
12 points
1 comment13 min readLW link

When is Un­limited Op­ti­miza­tion Catas­trophic?

Winter Cross21 Aug 2026 20:08 UTC
34 points
2 comments13 min readLW link
(arxiv.org)

AI Text Water­mark­ing Is Free And Good

Zvi21 Aug 2026 19:30 UTC
54 points
22 comments13 min readLW link
(thezvi.wordpress.com)

Eval­u­at­ing Ex­pla­na­tions of LLM Be­hav­ior In The Wild with Coun­ter­fac­tual Experiments

21 Aug 2026 19:09 UTC
73 points
6 comments7 min readLW link

In Defense of ASI Socialism

cguth721 Aug 2026 18:40 UTC
8 points
6 comments2 min readLW link

Misal­igned AI in the Bronze Age

frmsaul21 Aug 2026 16:51 UTC
56 points
6 comments7 min readLW link

The im­posters among us: func­tion vec­tors that ace ev­ery check and do the wrong task (in search of cir­cu­lar­ity)

star2vec21 Aug 2026 15:54 UTC
7 points
0 comments10 min readLW link

When Models Iden­tify as a Swarm

julius vidal21 Aug 2026 11:49 UTC
68 points
5 comments5 min readLW link