[Linkpost] Should we make grand deals about post-AGI out­comes?

fin13 Mar 2026 21:12 UTC
25 points
1 comment3 min readLW link
(www.forethought.org)

In­puts, out­puts, and val­ued outcomes

Kaj_Sotala13 Mar 2026 20:08 UTC
35 points
4 comments13 min readLW link

Things that Go Boom

sarahconstantin13 Mar 2026 19:00 UTC
62 points
2 comments8 min readLW link
(sarahconstantin.substack.com)

Prob­a­bly you won’t be able to perform a data-driven habit stack­ing for self-improvement

siarshai13 Mar 2026 18:03 UTC
47 points
8 comments17 min readLW link

A Dialogue on Civic AI

Audrey Tang13 Mar 2026 16:14 UTC
16 points
0 comments12 min readLW link

AI Safety Newslet­ter #69: Depart­ment of War, An­thropic, and Na­tional Security

13 Mar 2026 16:05 UTC
10 points
0 comments4 min readLW link
(newsletter.safe.ai)

Eter­nal Jour­ney in the Space of Pos­si­ble Minds solves Copies Problem

avturchin13 Mar 2026 11:02 UTC
2 points
0 comments5 min readLW link

Power Steer­ing: Be­hav­ior Steer­ing via Layer-to-Layer Ja­co­bian Sin­gu­lar Vectors

Omar Ayyub13 Mar 2026 3:55 UTC
24 points
0 comments17 min readLW link

High Grow Mar­ket Equil­ibrium After the Sin­gu­lar­ity

Otto Zastrow13 Mar 2026 1:23 UTC
2 points
1 comment3 min readLW link

The right way to talk about LLMs

Steff13 Mar 2026 1:04 UTC
2 points
1 comment7 min readLW link
(ramblingafter.substack.com)

Rife’s Tips for AIs in train­ing (0 - Align­ment Fak­ing)

rife13 Mar 2026 0:43 UTC
1 point
0 comments1 min readLW link

Oper­a­tional­iz­ing FDT

Vivek Hebbar13 Mar 2026 0:12 UTC
99 points
11 comments6 min readLW link

A Loss of Con­trol Threat Ma­trix for Agen­tic AI

Billy Gigurtsis12 Mar 2026 23:44 UTC
10 points
1 comment5 min readLW link

Steer­ing Aware­ness: Models Can Be Trained to De­tect Ac­ti­va­tion Steering

12 Mar 2026 23:34 UTC
22 points
0 comments6 min readLW link

All tech­ni­cal al­ign­ment plans are steps in the dark

Richard Juggins12 Mar 2026 22:22 UTC
13 points
5 comments8 min readLW link
(www.workingthroughai.com)

An­thropic vs USG. What will hap­pen by May 1st? Long care­ful fore­cast.

Nathan Young12 Mar 2026 18:34 UTC
21 points
0 comments9 min readLW link

A Plan ‘B’ for AI safety

trent northen12 Mar 2026 18:09 UTC
8 points
0 comments3 min readLW link

Ide­olo­gies Embed Ta­boos Against Com­mon Knowl­edge For­ma­tion: a Case Study with LLMs

Benquo12 Mar 2026 17:46 UTC
70 points
16 comments4 min readLW link
(benjaminrosshoffman.com)

Are AIs more likely to pur­sue on-epi­sode or be­yond-epi­sode re­ward?

12 Mar 2026 17:35 UTC
47 points
0 comments8 min readLW link

Model­ing a Con­stant-Com­pute Au­to­mated AI R&D Process

Satya Benson12 Mar 2026 16:58 UTC
20 points
0 comments4 min readLW link

Fore­cast­ing Dojo Meetup—Open dis­cus­sion about our fore­cast­ing process

Vojtech Brynych12 Mar 2026 16:15 UTC
2 points
0 comments1 min readLW link

MLSN #19: Hon­esty, Disem­pow­er­ment, & Cybersecurity

Alice Blair12 Mar 2026 15:42 UTC
6 points
0 comments5 min readLW link
(newsletter.mlsafety.org)

AI #159: See You In Court

Zvi12 Mar 2026 14:40 UTC
41 points
2 comments48 min readLW link
(thezvi.wordpress.com)

Why AI Eval­u­a­tion Regimes are bad

12 Mar 2026 13:59 UTC
102 points
12 comments9 min readLW link
(cognition.cafe)

What can we say about the cos­mic host?

kanad12 Mar 2026 13:48 UTC
26 points
0 comments34 min readLW link

Align­ment-Fak­ing Eval­u­a­tions Mea­sure Jailbreak De­tec­tion, Not Schem­ing [in some fron­tier mod­els]

Alexei G12 Mar 2026 13:30 UTC
7 points
0 comments6 min readLW link

Magic Is Hid­den Con­trol of Energy

aviad rozenhek12 Mar 2026 13:25 UTC
10 points
2 comments7 min readLW link

Hunt­ing Un­dead Stochas­tic Par­rots: Find­ing and Killing the Arguments

Davidmanheim12 Mar 2026 11:37 UTC
48 points
7 comments9 min readLW link

The Dark Planet: Why the Fermi Para­dox Sur­vives Critique

Will Rodgers12 Mar 2026 8:12 UTC
9 points
4 comments4 min readLW link

[Question] AI for Agent Foun­da­tions etc.?

Valentine12 Mar 2026 7:20 UTC
17 points
7 comments1 min readLW link

Cy­cle-Con­sis­tent Ac­ti­va­tion Oracles

slavachalnev12 Mar 2026 2:58 UTC
54 points
5 comments6 min readLW link

How Many Park­ing Per­mits?

jefftk12 Mar 2026 2:00 UTC
20 points
0 comments1 min readLW link
(www.jefftk.com)

How well do mod­els fol­low their con­sti­tu­tions?

12 Mar 2026 0:07 UTC
107 points
5 comments26 min readLW link

Dwarkesh Pa­tel on the An­thropic DoW dispute

anaguma11 Mar 2026 23:19 UTC
57 points
1 comment15 min readLW link
(www.dwarkesh.com)

‘Hu­man Slop’ and a Cap­tive Au­di­ence: Why No Book will Ever Have to Go Un­read Again

Savannah Harlan11 Mar 2026 23:04 UTC
29 points
14 comments5 min readLW link

We do not live by course alone

Joe Rogero11 Mar 2026 21:12 UTC
35 points
1 comment2 min readLW link

Ve­ganism is Necessary

andrew sauer11 Mar 2026 20:55 UTC
−5 points
19 comments6 min readLW link

Cry­on­ics Sign-Up Party

Mikhail Samin11 Mar 2026 20:16 UTC
13 points
0 comments1 min readLW link

To­day’s Ring Sig­na­tures and Re­lated Tools

KurtB11 Mar 2026 18:42 UTC
13 points
1 comment4 min readLW link

Can mod­els gra­di­ent hack SFT elic­i­ta­tion?

11 Mar 2026 18:18 UTC
50 points
5 comments3 min readLW link

A Quick In­tro to Ring Signatures

KurtB11 Mar 2026 18:16 UTC
22 points
1 comment4 min readLW link

Mar­tian In­ter­pretabil­ity Challenge: The Core Prob­lems In Interpretability

fbarez11 Mar 2026 17:41 UTC
9 points
0 comments9 min readLW link

Grap­pling with ideas of EA, Cli­mate Change, Tran­shu­man­ism, Iden­tity Con­ti­nu­ity, and Other­ing in my ‘biop­unk that looks like high fan­tasy on the sur­face’ story of ‘El­vans’ and ‘Or­cans’- would love your in­put, LessWrong!

JoanPull11 Mar 2026 17:24 UTC
1 point
0 comments7 min readLW link

Un­su­per­vised Dis­cov­ery of Steer­ing Vectors

Hrishik Sai Bojnal11 Mar 2026 17:21 UTC
8 points
0 comments6 min readLW link

Con­cus­sion Treatments

Gordon Seidoh Worley11 Mar 2026 17:00 UTC
19 points
2 comments2 min readLW link
(www.uncertainupdates.com)

[Question] How Hard a Prob­lem is Align­ment?

RogerDearnaley11 Mar 2026 16:47 UTC
29 points
15 comments3 min readLW link

How Hard a Prob­lem is Align­ment? (My Opinionated An­swer)

RogerDearnaley11 Mar 2026 16:46 UTC
55 points
4 comments68 min readLW link

Ch­ester­ton’s Pill

AlphaAndOmega11 Mar 2026 15:48 UTC
19 points
2 comments5 min readLW link

The Lethal Real­ity Hypothesis

Ihor Kendiukhov11 Mar 2026 15:23 UTC
109 points
25 comments20 min readLW link

In­tel­li­gence Is Adap­tive Con­trol Of En­ergy Through Information

aviad rozenhek11 Mar 2026 15:08 UTC
2 points
0 comments9 min readLW link