Rea­son­ing Models Strug­gle to Con­trol Their Chains of Thought

5 Mar 2026 22:37 UTC
76 points
9 comments3 min readLW link

Per­son­al­ity Self-Replicators

eggsyntax5 Mar 2026 20:30 UTC
173 points
51 comments10 min readLW link

Salient Direc­tions in AI Control

Bruce W. Lee5 Mar 2026 19:38 UTC
13 points
0 comments14 min readLW link
(brucewlee.com)

Models have lin­ear rep­re­sen­ta­tions of what tasks they like

OscarGilg5 Mar 2026 18:44 UTC
55 points
16 comments11 min readLW link

AI Safety Has 12 Months Left

mhdempsey5 Mar 2026 16:37 UTC
40 points
10 comments6 min readLW link
(mhdempsey.substack.com)

Have Amer­i­cans Be­come Less Violent Since 1980?

Benquo5 Mar 2026 16:11 UTC
78 points
6 comments14 min readLW link
(benjaminrosshoffman.com)

AI #158: The Depart­ment of War

Zvi5 Mar 2026 16:10 UTC
77 points
2 comments51 min readLW link
(thezvi.wordpress.com)

In­ves­ti­gat­ing Self-Fulfilling Misal­ign­ment and Col­lu­sion in AI Control

5 Mar 2026 15:05 UTC
15 points
0 comments5 min readLW link

Com­pu­ta­tion, Chess, and Lan­guage in Ar­tifi­cial In­tel­li­gence

Bill Benzon5 Mar 2026 12:57 UTC
6 points
0 comments3 min readLW link

Vibe Cod­ing crip­ples the mind

spookyuser5 Mar 2026 10:29 UTC
−12 points
4 comments4 min readLW link

Ra­tional Chess

8495 Mar 2026 9:57 UTC
5 points
19 comments2 min readLW link

A Be­havi­oural and Rep­re­sen­ta­tional Eval­u­a­tion of Goal-di­rect­ed­ness in Lan­guage Model Agents

5 Mar 2026 1:08 UTC
20 points
0 comments7 min readLW link

Fe­bru­ary 2026 Links

nomagicpill5 Mar 2026 1:03 UTC
6 points
2 comments7 min readLW link
(nomagicpill.substack.com)

Text Com­pres­sion Can Help Se­cure Model Weights

Roy Rinberg4 Mar 2026 23:30 UTC
45 points
12 comments10 min readLW link

A sum­mary of Con­den­sa­tion and its re­la­tion to Nat­u­ral Latents

4 Mar 2026 22:22 UTC
86 points
0 comments10 min readLW link

Maybe there’s a pat­tern here?

dynomight4 Mar 2026 20:32 UTC
193 points
45 comments7 min readLW link

Gem­ini 3.1 Pro Aces Bench­marks, I Suppose

Zvi4 Mar 2026 20:10 UTC
27 points
2 comments9 min readLW link
(thezvi.wordpress.com)

Gen Z and AI: Ed­u­ca­tion, Well-be­ing, and the Labour Mar­ket

Jason Hung4 Mar 2026 19:19 UTC
9 points
0 comments11 min readLW link

Is GDP a Kind of Fac­tory?

Benquo4 Mar 2026 18:00 UTC
59 points
3 comments10 min readLW link
(benjaminrosshoffman.com)

Make Pow­er­ful Machines Verifiable

Naci Cankaya4 Mar 2026 14:20 UTC
22 points
4 comments4 min readLW link

How a Pinky Promise once stopped a war in the Mid­dle East.

positivesum4 Mar 2026 13:30 UTC
29 points
1 comment4 min readLW link

Split Per­son­al­ity Train­ing can de­tect Align­ment Faking

Florian_Dietz4 Mar 2026 11:49 UTC
39 points
0 comments6 min readLW link

Physics of RL: Toy scal­ing laws for the emer­gence of re­ward-seeking

Alex Meinke4 Mar 2026 8:12 UTC
120 points
9 comments10 min readLW link

Sa­cred val­ues of fu­ture AIs

Cleo Nardo4 Mar 2026 7:47 UTC
58 points
4 comments5 min readLW link

OpenAI’s surveillance lan­guage has many po­ten­tial loop­holes and they can do better

Tom Smith4 Mar 2026 4:25 UTC
153 points
4 comments10 min readLW link

Mass surveillance, red lines, and a crazy weekend

Boaz Barak4 Mar 2026 4:24 UTC
35 points
26 comments5 min readLW link

Lie To Me, But At Least Don’t Bullshit

Czynski4 Mar 2026 2:20 UTC
20 points
4 comments5 min readLW link
(dangeroussincerity.substack.com)

[Question] LLM co­her­en­ti­za­tion as an ob­vi­ous low-hang­ing fruit to try?

Épiphanie Gédéon4 Mar 2026 0:59 UTC
26 points
2 comments2 min readLW link

Milder tem­per­a­ture makes a hell stable

Joachim Bartosik3 Mar 2026 22:25 UTC
17 points
1 comment2 min readLW link

Mass Surveillance w/​ LLMs is the De­fault Out­come. Con­tracts Won’t Change That.

Logan Riggs3 Mar 2026 21:18 UTC
43 points
1 comment2 min readLW link

A Tale of Three Contracts

Zvi3 Mar 2026 20:30 UTC
45 points
3 comments16 min readLW link
(thezvi.wordpress.com)

An Align­ment Jour­nal: Com­ing Soon

3 Mar 2026 20:27 UTC
271 points
34 comments6 min readLW link
(blog.alignmentjournal.org)

Cur­rent ac­ti­va­tion or­a­cles are hard to use

3 Mar 2026 19:33 UTC
83 points
4 comments16 min readLW link

[Question] Ques­tion: Why is the goal of AI safety not ‘moral ma­chines’?

Mordechai Rorvig3 Mar 2026 18:16 UTC
10 points
15 comments1 min readLW link

An Age Of Promethean Ambitions

sonicrocketman3 Mar 2026 18:09 UTC
0 points
2 comments4 min readLW link
(brianschrader.com)

White-Box At­tacks on the Best Open-Weight Model: CCP Bias vs. Safety Train­ing in Kimi K2.5

Corm3 Mar 2026 17:47 UTC
16 points
2 comments5 min readLW link

I Had Claude Read Every AI Safety Paper Since 2020, Here’s the DB

Corm3 Mar 2026 17:47 UTC
59 points
13 comments3 min readLW link

Con­sti­tu­tional Black-Box Mon­i­tor­ing for Schem­ing in LLM Agents

3 Mar 2026 17:01 UTC
28 points
0 comments2 min readLW link

LASR Labs Sum­mer 2026 ap­pli­ca­tions are open!

3 Mar 2026 15:42 UTC
25 points
0 comments3 min readLW link

AI com­pa­nies and the 99% lethal au­tonomous weapons myth

User_Luke3 Mar 2026 12:18 UTC
−1 points
1 comment2 min readLW link

I’m con­fused by the change in the METR trend

Expertium3 Mar 2026 11:30 UTC
46 points
17 comments2 min readLW link

Zurich AI Safety is hiring a Director

3 Mar 2026 10:29 UTC
21 points
0 comments3 min readLW link

Game Rec­og­nizes Game

eva_3 Mar 2026 10:09 UTC
105 points
15 comments12 min readLW link

Mon­day AI Radar #15

Against Moloch3 Mar 2026 5:23 UTC
13 points
1 comment7 min readLW link

Me­mory De­cod­ing Jour­nal Club: En­gram cell con­nec­tivity as a mechanism for in­for­ma­tion en­cod­ing and mem­ory func­tion

Devin Ward3 Mar 2026 1:32 UTC
3 points
0 comments1 min readLW link

In-con­text learn­ing of rep­re­sen­ta­tions can be ex­plained by in­duc­tion circuits

Andy Arditi2 Mar 2026 23:58 UTC
51 points
0 comments9 min readLW link
(iclr-blogposts.github.io)

Sin­gle Direc­tion vs Low-Rank Re­fusal in Small LLMs

IvanC2 Mar 2026 23:14 UTC
12 points
0 comments8 min readLW link

Be­ing am­bi­tious in soulful altruism

pandamonium2 Mar 2026 21:15 UTC
5 points
0 comments4 min readLW link

Notes on the “Heart of Dark­ness”

dominicq2 Mar 2026 20:11 UTC
5 points
0 comments4 min readLW link
(sundaystopwatch.eu)

[Question] Can LLM chat be less pro­lix?

jbash2 Mar 2026 19:54 UTC
21 points
9 comments2 min readLW link