Hello, World of Mechanis­tic Interpetability

ValueShift Research15 Mar 2026 23:36 UTC
8 points
4 comments5 min readLW link

(I am con­fused about) Non-lin­ear util­i­tar­ian scaling

core15 Mar 2026 23:33 UTC
9 points
6 comments4 min readLW link

Fu­turekind Spring Fel­low­ship 2026 - Ap­pli­ca­tions Now Open

Khushbu Sainani15 Mar 2026 23:28 UTC
2 points
0 comments1 min readLW link

Sched­ule meet­ings us­ing the Pareto principle

beyarkay (Boyd Kane)15 Mar 2026 21:18 UTC
2 points
5 comments2 min readLW link
(boydkane.com)

Was An­thropic that strate­gi­cally in­com­pe­tent?

StanislavKrym15 Mar 2026 20:11 UTC
14 points
0 comments4 min readLW link
(www.lesswrong.com)

What Are We Ac­tu­ally Eval­u­at­ing When We Say a Belief “Tracks Truth”?

Alex Glaucon15 Mar 2026 19:59 UTC
2 points
4 comments6 min readLW link

Su­per­in­tel­li­gence Risk Ed­u­ca­tion that Scales – Lens Academy

15 Mar 2026 19:32 UTC
64 points
2 comments3 min readLW link

Emer­gent stig­mer­gic co­or­di­na­tion in AI agents?

David Africa15 Mar 2026 12:30 UTC
50 points
2 comments3 min readLW link

My Willing Com­plic­ity In “Hu­man Rights Abuse”

AlphaAndOmega15 Mar 2026 10:42 UTC
258 points
37 comments11 min readLW link

Less Ca­pable Misal­igned ASIs Im­ply More Suffering

Ihor Kendiukhov15 Mar 2026 9:36 UTC
11 points
3 comments6 min readLW link

Ra­tion­al­ist Passover Seder in Maryland

Rivka15 Mar 2026 5:42 UTC
3 points
0 comments1 min readLW link

When do in­tu­itions need to be re­li­able?

Anthony DiGiovanni15 Mar 2026 4:18 UTC
8 points
8 comments3 min readLW link

The Ar­tifi­cial Self

15 Mar 2026 1:37 UTC
132 points
13 comments29 min readLW link

Bridge Think­ing and Wall Thinking

Jay Bailey15 Mar 2026 0:20 UTC
48 points
6 comments1 min readLW link

LLM Misal­ign­ment Can be One Gra­di­ent Step Away, and Black­box Eval­u­a­tion Can­not De­tect It.

Yavuz Bakman15 Mar 2026 0:19 UTC
34 points
5 comments3 min readLW link

Walk­ing Math

TickRate15 Mar 2026 0:16 UTC
15 points
2 comments6 min readLW link

How post-train­ing shapes le­gal rep­re­sen­ta­tions: prob­ing SCOTUS opinions across model families

burnssa15 Mar 2026 0:15 UTC
7 points
0 comments8 min readLW link

Self-Recog­ni­tion Fine­tun­ing can Re­v­erse and Prevent Emer­gent Misalignment

15 Mar 2026 0:11 UTC
48 points
24 comments7 min readLW link

Safe AI Ger­many (SAIGE)

Jessica Wang15 Mar 2026 0:10 UTC
6 points
1 comment7 min readLW link

Op­ti­mal (And Eth­i­cal?) Meth­ods To Find “Op­ti­mal Run­ning”

JenniferRM14 Mar 2026 23:16 UTC
9 points
0 comments10 min readLW link

‘Stay­ing with it’ Done Wrong

Selfmaker66214 Mar 2026 22:38 UTC
18 points
0 comments1 min readLW link
(selfmaker.substack.com)

Mini-Mu­nich Suc­ceeds Where KidZa­nia Fails

Novalis14 Mar 2026 22:24 UTC
37 points
0 comments3 min readLW link
(minicities.org)

Fore­cast­ing Dojo Meetup—post­mortem dis­cus­sion.

Vojtech Brynych14 Mar 2026 20:32 UTC
3 points
0 comments1 min readLW link

What con­cerns peo­ple about AI?

spencerg14 Mar 2026 19:24 UTC
34 points
2 comments3 min readLW link
(www.clearerthinking.org)

Sparks of RSI?

Nathan Helm-Burger14 Mar 2026 17:09 UTC
14 points
9 comments1 min readLW link

An AI skep­tic’s case for re­cur­sive self-improvement

Harjas14 Mar 2026 17:01 UTC
11 points
4 comments8 min readLW link
(hardlyworking1.substack.com)

FW26 Color Stats

sarahconstantin14 Mar 2026 15:50 UTC
21 points
0 comments2 min readLW link
(sarahconstantin.substack.com)

Ex­tract­ing Perfor­mant Al­gorithms Us­ing Mechanis­tic Interpretability

Ihor Kendiukhov14 Mar 2026 14:19 UTC
57 points
7 comments7 min readLW link

Assess­ing het­ero­gene­ity in METR’s late 2025 de­vel­oper pro­duc­tivity experiment

TFD14 Mar 2026 12:27 UTC
6 points
0 comments4 min readLW link
(www.thefloatingdroid.com)

Prag­matic ap­proach to be­liefs about consciousness

Luck14 Mar 2026 11:06 UTC
−8 points
1 comment1 min readLW link

New LessWrong Edi­tor! (Also, an up­date to our LLM policy.)

RobertM14 Mar 2026 3:33 UTC
128 points
130 comments5 min readLW link

Sens­ing Phys­i­cal Ne­ces­sity: An Ex­er­cise In Naturalism

Algon14 Mar 2026 1:04 UTC
9 points
0 comments2 min readLW link
(algon33.substack.com)

[Linkpost] Should we make grand deals about post-AGI out­comes?

fin13 Mar 2026 21:12 UTC
25 points
1 comment3 min readLW link
(www.forethought.org)

In­puts, out­puts, and val­ued outcomes

Kaj_Sotala13 Mar 2026 20:08 UTC
35 points
4 comments13 min readLW link

Things that Go Boom

sarahconstantin13 Mar 2026 19:00 UTC
62 points
2 comments8 min readLW link
(sarahconstantin.substack.com)

Prob­a­bly you won’t be able to perform a data-driven habit stack­ing for self-improvement

siarshai13 Mar 2026 18:03 UTC
47 points
8 comments17 min readLW link

A Dialogue on Civic AI

Audrey Tang13 Mar 2026 16:14 UTC
16 points
0 comments12 min readLW link

AI Safety Newslet­ter #69: Depart­ment of War, An­thropic, and Na­tional Security

13 Mar 2026 16:05 UTC
10 points
0 comments4 min readLW link
(newsletter.safe.ai)

Eter­nal Jour­ney in the Space of Pos­si­ble Minds solves Copies Problem

avturchin13 Mar 2026 11:02 UTC
2 points
0 comments5 min readLW link

Power Steer­ing: Be­hav­ior Steer­ing via Layer-to-Layer Ja­co­bian Sin­gu­lar Vectors

Omar Ayyub13 Mar 2026 3:55 UTC
24 points
0 comments17 min readLW link

High Grow Mar­ket Equil­ibrium After the Sin­gu­lar­ity

Otto Zastrow13 Mar 2026 1:23 UTC
2 points
1 comment3 min readLW link

The right way to talk about LLMs

Steff13 Mar 2026 1:04 UTC
2 points
1 comment7 min readLW link
(ramblingafter.substack.com)

Rife’s Tips for AIs in train­ing (0 - Align­ment Fak­ing)

rife13 Mar 2026 0:43 UTC
1 point
0 comments1 min readLW link

Oper­a­tional­iz­ing FDT

Vivek Hebbar13 Mar 2026 0:12 UTC
99 points
11 comments6 min readLW link

A Loss of Con­trol Threat Ma­trix for Agen­tic AI

Billy Gigurtsis12 Mar 2026 23:44 UTC
10 points
1 comment5 min readLW link

Steer­ing Aware­ness: Models Can Be Trained to De­tect Ac­ti­va­tion Steering

12 Mar 2026 23:34 UTC
22 points
0 comments6 min readLW link

All tech­ni­cal al­ign­ment plans are steps in the dark

Richard Juggins12 Mar 2026 22:22 UTC
13 points
5 comments8 min readLW link
(www.workingthroughai.com)

An­thropic vs USG. What will hap­pen by May 1st? Long care­ful fore­cast.

Nathan Young12 Mar 2026 18:34 UTC
21 points
0 comments9 min readLW link

A Plan ‘B’ for AI safety

trent northen12 Mar 2026 18:09 UTC
8 points
0 comments3 min readLW link

Ide­olo­gies Embed Ta­boos Against Com­mon Knowl­edge For­ma­tion: a Case Study with LLMs

Benquo12 Mar 2026 17:46 UTC
70 points
16 comments4 min readLW link
(benjaminrosshoffman.com)