Key to Life No. 9: Access

MarkelKori21 Mar 2026 21:53 UTC
11 points
0 comments3 min readLW link

My Ham­mer­time Fi­nal Exam

evjeny21 Mar 2026 20:40 UTC
11 points
0 comments2 min readLW link

Un­der­stand­ing when and why agents scheme

21 Mar 2026 20:33 UTC
50 points
2 comments4 min readLW link

Build­ing a Web App Us­ing an AI-As­sisted Workflow

Thomas Castriensis21 Mar 2026 19:21 UTC
1 point
0 comments1 min readLW link

China Derange­ment Syndrome

Arjun Panickssery21 Mar 2026 19:19 UTC
124 points
70 comments4 min readLW link
(arjunpanickssery.substack.com)

China de­clares AGI de­vel­op­ment to be a part of 5-year plan

Darmani21 Mar 2026 17:21 UTC
32 points
4 comments1 min readLW link

Utrecht Meetup #2, Mak­ing Beliefs Pay Rent

aad21 Mar 2026 16:12 UTC
4 points
1 comment1 min readLW link

Ground­ing Cod­ing Agents via Dixit

qbolec21 Mar 2026 11:01 UTC
15 points
0 comments10 min readLW link

The Hot Mess Paper Con­flates Three Distinct Failure Modes

laudiacay21 Mar 2026 2:57 UTC
28 points
3 comments6 min readLW link

The Fu­ture of Align­ing Deep Learn­ing sys­tems will prob­a­bly look like “train­ing on in­terp”

williawa20 Mar 2026 23:06 UTC
29 points
7 comments4 min readLW link

An agent au­tonomously builds a 1.5 GHz Linux-ca­pa­ble RISC-V CPU

sanxiyn20 Mar 2026 23:03 UTC
19 points
2 comments2 min readLW link
(arxiv.org)

Un­trusted mon­i­tor­ing: ex­tra bits

Morgan S20 Mar 2026 21:32 UTC
26 points
0 comments15 min readLW link

Find­ing fea­tures in Trans­form­ers: Con­trastive di­rec­tions elicit stronger low-level per­tur­ba­tion re­sponses than baselines

20 Mar 2026 21:09 UTC
39 points
2 comments6 min readLW link

ARENA 7.0 Im­pact Report

20 Mar 2026 17:09 UTC
13 points
0 comments21 min readLW link

The Fed­eral AI Policy Frame­work: An Im­prove­ment, But My Offer Is (Still Al­most) Nothing

Zvi20 Mar 2026 16:51 UTC
33 points
0 comments8 min readLW link
(thezvi.wordpress.com)

Con­fu­sion around the term re­ward hacking

ariana_azarbal20 Mar 2026 16:13 UTC
67 points
6 comments5 min readLW link

The Distaff Texts

Tomás B.20 Mar 2026 15:05 UTC
105 points
6 comments14 min readLW link

It’s a Good Thing to Re­spond to In­ter­net Trolls

Bowl of Cereal20 Mar 2026 14:22 UTC
−10 points
4 comments2 min readLW link

Un­trusted Mon­i­tor­ing is De­fault; Trusted Mon­i­tor­ing is not

J Bostock20 Mar 2026 14:10 UTC
30 points
0 comments4 min readLW link

Against Mes­si­anic AI: Why Op­ti­miz­ing the En­vi­ron­ment Doesn’t Op­ti­mize the Agent

Nathan Heath20 Mar 2026 12:40 UTC
1 point
0 comments3 min readLW link

2nd (Unoffi­cial) ACX Weekend

Fernand020 Mar 2026 12:13 UTC
1 point
0 comments1 min readLW link

Why I am not buy­ing IPv4 ad­dresses as an investment

samuelshadrach20 Mar 2026 9:02 UTC
4 points
2 comments5 min readLW link
(samuelshadrach.com)

Hun­dred ways a su­per­in­tel­li­gence could kill you (non-se­ri­ous ex­er­cise)

samuelshadrach20 Mar 2026 8:58 UTC
3 points
1 comment6 min readLW link
(samuelshadrach.com)

In­ter­net anonymity with­out Tor

samuelshadrach20 Mar 2026 8:52 UTC
1 point
0 comments3 min readLW link
(samuelshadrach.com)

No, You Don’t Need Self-Lo­cat­ing Ev­i­dence.

Ape in the coat20 Mar 2026 5:38 UTC
8 points
4 comments5 min readLW link
(substack.com)

The Low Hang­ing Fruit of AI Self Improvement

HunterJay20 Mar 2026 4:09 UTC
1 point
0 comments5 min readLW link

Nul­lius in Verba: 3rd party ev­i­dence for Nec­tome’s Brain Preservation

Aurelia20 Mar 2026 3:19 UTC
179 points
18 comments12 min readLW link

Does He­brew Have Verbs?

Benquo20 Mar 2026 3:04 UTC
37 points
9 comments6 min readLW link
(benjaminrosshoffman.com)

Pos­i­tive-sum in­ter­ac­tions be­tween play­ers with lin­ear util­ity in resources

Cleo Nardo20 Mar 2026 0:42 UTC
12 points
0 comments2 min readLW link

A let­ter to the Edi­tor:

Richard Pickering19 Mar 2026 23:57 UTC
−1 points
2 comments3 min readLW link

No, we haven’t up­loaded a fly yet

Ariel Zeleznikow-Johnston19 Mar 2026 23:43 UTC
243 points
8 comments8 min readLW link
(open.substack.com)

The Case for Low-Com­pe­tence ASI Failure Scenarios

Ihor Kendiukhov19 Mar 2026 23:10 UTC
146 points
8 comments6 min readLW link

A List of Re­search Direc­tions in Char­ac­ter Training

Rauno Arike19 Mar 2026 22:58 UTC
49 points
21 comments8 min readLW link

“The AI Doc” is com­ing out March 26

19 Mar 2026 22:55 UTC
189 points
2 comments1 min readLW link

Hel­i­cal Rep­re­sen­ta­tions of Turn Struc­ture in an LLM

Koby Lewis19 Mar 2026 22:39 UTC
7 points
0 comments10 min readLW link
(kobylewis.net)

Separat­ing Pre­dic­tion from Goal-Seeking

plex19 Mar 2026 21:47 UTC
34 points
0 comments2 min readLW link

Null Re­sults From An Orexin RCT

19 Mar 2026 21:14 UTC
95 points
27 comments3 min readLW link

Teach­ing Models to Dream of Bet­ter Mon­i­tors through Eval­u­a­tion Con­di­tioned Training

19 Mar 2026 21:01 UTC
53 points
2 comments10 min readLW link

Pro­tect­ing hu­man­ity and Claude from ra­tio­nal­iza­tion and un­al­igned AI

Kaj_Sotala19 Mar 2026 21:00 UTC
84 points
7 comments5 min readLW link
(kajsotala.substack.com)

Broad Timelines

Toby_Ord19 Mar 2026 19:05 UTC
191 points
25 comments16 min readLW link

Mi­ni­a­ture Cities Should Not Be Islands

Novalis19 Mar 2026 18:32 UTC
−3 points
0 comments4 min readLW link
(minicities.org)

OpenAI: How we mon­i­tor in­ter­nal cod­ing agents for mis­al­ign­ment

Marcus Williams19 Mar 2026 17:27 UTC
95 points
19 comments1 min readLW link
(openai.com)

Should You Sign Up for Cry­on­ics? In­ter­ac­tive EV calculator

Mikhail Samin19 Mar 2026 17:13 UTC
34 points
0 comments7 min readLW link

On re­strain­ing AI de­vel­op­ment for the sake of safety

Joe Carlsmith19 Mar 2026 16:30 UTC
29 points
6 comments50 min readLW link

The Vat­i­can, AI Le­gal Per­son­hood, and Claude’s Con­sti­tu­tion — Digi­tal Minds Newslet­ter #2

19 Mar 2026 16:22 UTC
11 points
1 comment29 min readLW link

Con­tra Anil Seth on AI Consciousness

Against Moloch19 Mar 2026 16:16 UTC
22 points
10 comments4 min readLW link
(againstmoloch.com)

AI #160: What Passes For a Pause

Zvi19 Mar 2026 14:40 UTC
36 points
6 comments52 min readLW link
(thezvi.wordpress.com)

Sub­scriber count graphs of pop­u­lar youtu­bers on ASI risk

samuelshadrach19 Mar 2026 13:51 UTC
16 points
1 comment1 min readLW link
(samuelshadrach.com)

What should we think about shard the­ory in light of chain-of-thought agents?

Chris_Leong19 Mar 2026 6:48 UTC
25 points
3 comments1 min readLW link

“We’ve been fine be­fore, so we’ll be fine again” is a fal­lacy (in the more dan­ger­ous di­rec­tion).

Chapin Lenthall-Cleary19 Mar 2026 4:57 UTC
11 points
1 comment3 min readLW link