Dwarkesh Pa­tel on the An­thropic DoW dispute

anaguma11 Mar 2026 23:19 UTC
57 points
1 comment15 min readLW link
(www.dwarkesh.com)

‘Hu­man Slop’ and a Cap­tive Au­di­ence: Why No Book will Ever Have to Go Un­read Again

Savannah Harlan11 Mar 2026 23:04 UTC
29 points
14 comments5 min readLW link

We do not live by course alone

Joe Rogero11 Mar 2026 21:12 UTC
35 points
1 comment2 min readLW link

Ve­ganism is Necessary

andrew sauer11 Mar 2026 20:55 UTC
−5 points
19 comments6 min readLW link

Cry­on­ics Sign-Up Party

Mikhail Samin11 Mar 2026 20:16 UTC
13 points
0 comments1 min readLW link

To­day’s Ring Sig­na­tures and Re­lated Tools

KurtB11 Mar 2026 18:42 UTC
13 points
1 comment4 min readLW link

Can mod­els gra­di­ent hack SFT elic­i­ta­tion?

11 Mar 2026 18:18 UTC
50 points
5 comments3 min readLW link

A Quick In­tro to Ring Signatures

KurtB11 Mar 2026 18:16 UTC
22 points
1 comment4 min readLW link

Mar­tian In­ter­pretabil­ity Challenge: The Core Prob­lems In Interpretability

fbarez11 Mar 2026 17:41 UTC
9 points
0 comments9 min readLW link

Grap­pling with ideas of EA, Cli­mate Change, Tran­shu­man­ism, Iden­tity Con­ti­nu­ity, and Other­ing in my ‘biop­unk that looks like high fan­tasy on the sur­face’ story of ‘El­vans’ and ‘Or­cans’- would love your in­put, LessWrong!

JoanPull11 Mar 2026 17:24 UTC
1 point
0 comments7 min readLW link

Un­su­per­vised Dis­cov­ery of Steer­ing Vectors

Hrishik Sai Bojnal11 Mar 2026 17:21 UTC
8 points
0 comments6 min readLW link

Con­cus­sion Treatments

Gordon Seidoh Worley11 Mar 2026 17:00 UTC
19 points
2 comments2 min readLW link
(www.uncertainupdates.com)

[Question] How Hard a Prob­lem is Align­ment?

RogerDearnaley11 Mar 2026 16:47 UTC
29 points
15 comments3 min readLW link

How Hard a Prob­lem is Align­ment? (My Opinionated An­swer)

RogerDearnaley11 Mar 2026 16:46 UTC
55 points
4 comments68 min readLW link

Ch­ester­ton’s Pill

AlphaAndOmega11 Mar 2026 15:48 UTC
19 points
2 comments5 min readLW link

The Lethal Real­ity Hypothesis

Ihor Kendiukhov11 Mar 2026 15:23 UTC
109 points
25 comments20 min readLW link

In­tel­li­gence Is Adap­tive Con­trol Of En­ergy Through Information

aviad rozenhek11 Mar 2026 15:08 UTC
2 points
0 comments9 min readLW link

GPT-5.4 Is A Sub­stan­tial Upgrade

Zvi11 Mar 2026 14:00 UTC
23 points
5 comments24 min readLW link
(thezvi.wordpress.com)

The Refined Coun­ter­fac­tual Pri­soner’s Dilemma: An At­tempt to Ex­plode De­ci­sion-The­o­retic Con­se­quen­tial­ism

Chris_Leong11 Mar 2026 12:32 UTC
18 points
20 comments2 min readLW link

Helping Friends, Harm­ing Foes: Test­ing Trib­al­ism in Lan­guage Models

11 Mar 2026 12:06 UTC
10 points
0 comments9 min readLW link

AIs will be used in “un­hinged” configurations

Arthur Conmy11 Mar 2026 11:19 UTC
62 points
3 comments4 min readLW link

Neg­li­gent AI: Rea­son­able Care for AI Safety

Alex Mark11 Mar 2026 10:31 UTC
13 points
4 comments12 min readLW link

Less Dead

Aurelia11 Mar 2026 5:07 UTC
555 points
140 comments8 min readLW link

Con­flicted on Ramsey

jefftk11 Mar 2026 3:50 UTC
35 points
12 comments2 min readLW link
(www.jefftk.com)

Model weight preservation

tbs11 Mar 2026 1:53 UTC
3 points
0 comments21 min readLW link
(meditationsondigitalminds.substack.com)

San­ity Week­end Retrospective

10 Mar 2026 23:37 UTC
11 points
0 comments6 min readLW link

What do we know about AI com­pany em­ployee giv­ing?

David Scott Krueger10 Mar 2026 23:30 UTC
39 points
5 comments2 min readLW link

The Day After Move 37

Eneasz10 Mar 2026 23:05 UTC
65 points
2 comments6 min readLW link
(deathisbad.substack.com)

In­ter­view with Steven Byrnes on His Main­line Take­off Scenario

Liron10 Mar 2026 20:17 UTC
36 points
8 comments57 min readLW link
(doomdebates.com)

Au­ditBench: Eval­u­at­ing Align­ment Au­dit­ing Tech­niques on Models with Hid­den Behaviors

abhayesian10 Mar 2026 19:31 UTC
82 points
4 comments8 min readLW link
(alignment.anthropic.com)

Eco­nomic effi­ciency of­ten un­der­mines so­ciopoli­ti­cal autonomy

Richard_Ngo10 Mar 2026 19:30 UTC
166 points
38 comments12 min readLW link
(www.mindthefuture.info)

Let­ting Claude do Au­tonomous Re­search to Im­prove SAEs

chanind10 Mar 2026 18:52 UTC
102 points
16 comments7 min readLW link

Don’t Let LLMs Write For You

JustisMills10 Mar 2026 18:49 UTC
184 points
45 comments3 min readLW link
(justismills.substack.com)

Ques­tions to ask when ev­ery­one is shoot­ing them­selves in the foot

jasoncrawford10 Mar 2026 18:36 UTC
34 points
2 comments1 min readLW link

Harry Pot­ter Fan­fic­tion, which shows how the Dark Side reasons

Belle 10 Mar 2026 18:11 UTC
−3 points
0 comments3 min readLW link

The case for sa­ti­at­ing cheaply-satis­fied AI preferences

Alex Mallen10 Mar 2026 18:09 UTC
112 points
7 comments23 min readLW link

Gemma Needs Help

annajs10 Mar 2026 17:39 UTC
284 points
22 comments6 min readLW link

[Paper] When can we trust un­trusted mon­i­tor­ing? A safety case sketch across col­lu­sion strategies

10 Mar 2026 17:28 UTC
46 points
4 comments6 min readLW link

How Gen­er­a­tive AI Works: A Kalei­do­scope Metaphor

Disruptor10 Mar 2026 17:27 UTC
−9 points
0 comments11 min readLW link

Selec­tively re­duc­ing eval aware­ness and mur­der in Gemma 3 27B via steering

Matthias Murdych10 Mar 2026 17:26 UTC
8 points
0 comments3 min readLW link

Re­la­tional Harm: A Ther­a­pist’s Per­spec­tive on Con­ver­sa­tional AI

Asheley Korf10 Mar 2026 17:26 UTC
2 points
0 comments3 min readLW link

The Oper­a­tional Se­cu­rity Failure in An­thropic’s RSP v3

Ugurcan Arikan10 Mar 2026 17:08 UTC
2 points
0 comments4 min readLW link

Not Lov­ing Lik­ing What You See

Tomás B.10 Mar 2026 16:05 UTC
52 points
18 comments5 min readLW link

Load-Bear­ing Walls

sonicrocketman10 Mar 2026 14:29 UTC
3 points
2 comments1 min readLW link
(brianschrader.com)

Statis­ti­cism: How Cluster-Think­ing About Data Creates Blind Spots

Benquo10 Mar 2026 13:59 UTC
25 points
0 comments12 min readLW link
(benjaminrosshoffman.com)

Spon­ta­neous Sym­me­try Break­ing (Stat Mech Part 4)

J Bostock10 Mar 2026 13:21 UTC
16 points
2 comments4 min readLW link

Monthly Shorts 1/​22

KellerScholl10 Mar 2026 13:20 UTC
4 points
0 comments2 min readLW link
(keller.substack.com)

Why I don’t usu­ally recom­mend dead drops

samuelshadrach10 Mar 2026 13:13 UTC
4 points
2 comments4 min readLW link
(samuelshadrach.com)

Four Sce­nar­ios of Job-Re­duc­ing AI

KellerScholl10 Mar 2026 13:10 UTC
12 points
2 comments4 min readLW link
(keller.substack.com)

Un­der­stand­ing Rea­son­ing with Thought An­chors and Probes

10 Mar 2026 11:50 UTC
15 points
0 comments13 min readLW link