Text Com­pres­sion Can Help Se­cure Model Weights

Roy Rinberg4 Mar 2026 23:30 UTC
45 points
12 comments10 min readLW link

A sum­mary of Con­den­sa­tion and its re­la­tion to Nat­u­ral Latents

4 Mar 2026 22:22 UTC
86 points
0 comments10 min readLW link

Maybe there’s a pat­tern here?

dynomight4 Mar 2026 20:32 UTC
193 points
45 comments7 min readLW link

Gem­ini 3.1 Pro Aces Bench­marks, I Suppose

Zvi4 Mar 2026 20:10 UTC
27 points
2 comments9 min readLW link
(thezvi.wordpress.com)

Gen Z and AI: Ed­u­ca­tion, Well-be­ing, and the Labour Mar­ket

Jason Hung4 Mar 2026 19:19 UTC
9 points
0 comments11 min readLW link

Is GDP a Kind of Fac­tory?

Benquo4 Mar 2026 18:00 UTC
59 points
3 comments10 min readLW link
(benjaminrosshoffman.com)

Make Pow­er­ful Machines Verifiable

Naci Cankaya4 Mar 2026 14:20 UTC
22 points
4 comments4 min readLW link

How a Pinky Promise once stopped a war in the Mid­dle East.

positivesum4 Mar 2026 13:30 UTC
29 points
1 comment4 min readLW link

Split Per­son­al­ity Train­ing can de­tect Align­ment Faking

Florian_Dietz4 Mar 2026 11:49 UTC
39 points
0 comments6 min readLW link

Physics of RL: Toy scal­ing laws for the emer­gence of re­ward-seeking

Alex Meinke4 Mar 2026 8:12 UTC
120 points
9 comments10 min readLW link

Sa­cred val­ues of fu­ture AIs

Cleo Nardo4 Mar 2026 7:47 UTC
58 points
4 comments5 min readLW link

OpenAI’s surveillance lan­guage has many po­ten­tial loop­holes and they can do better

Tom Smith4 Mar 2026 4:25 UTC
153 points
4 comments10 min readLW link

Mass surveillance, red lines, and a crazy weekend

Boaz Barak4 Mar 2026 4:24 UTC
35 points
26 comments5 min readLW link

Lie To Me, But At Least Don’t Bullshit

Czynski4 Mar 2026 2:20 UTC
20 points
4 comments5 min readLW link
(dangeroussincerity.substack.com)

[Question] LLM co­her­en­ti­za­tion as an ob­vi­ous low-hang­ing fruit to try?

Épiphanie Gédéon4 Mar 2026 0:59 UTC
26 points
2 comments2 min readLW link

Milder tem­per­a­ture makes a hell stable

Joachim Bartosik3 Mar 2026 22:25 UTC
17 points
1 comment2 min readLW link

Mass Surveillance w/​ LLMs is the De­fault Out­come. Con­tracts Won’t Change That.

Logan Riggs3 Mar 2026 21:18 UTC
43 points
1 comment2 min readLW link

A Tale of Three Contracts

Zvi3 Mar 2026 20:30 UTC
45 points
3 comments16 min readLW link
(thezvi.wordpress.com)

An Align­ment Jour­nal: Com­ing Soon

3 Mar 2026 20:27 UTC
271 points
34 comments6 min readLW link
(blog.alignmentjournal.org)

Cur­rent ac­ti­va­tion or­a­cles are hard to use

3 Mar 2026 19:33 UTC
83 points
4 comments16 min readLW link

[Question] Ques­tion: Why is the goal of AI safety not ‘moral ma­chines’?

Mordechai Rorvig3 Mar 2026 18:16 UTC
10 points
15 comments1 min readLW link

An Age Of Promethean Ambitions

sonicrocketman3 Mar 2026 18:09 UTC
0 points
2 comments4 min readLW link
(brianschrader.com)

White-Box At­tacks on the Best Open-Weight Model: CCP Bias vs. Safety Train­ing in Kimi K2.5

Corm3 Mar 2026 17:47 UTC
16 points
2 comments5 min readLW link

I Had Claude Read Every AI Safety Paper Since 2020, Here’s the DB

Corm3 Mar 2026 17:47 UTC
59 points
13 comments3 min readLW link

Con­sti­tu­tional Black-Box Mon­i­tor­ing for Schem­ing in LLM Agents

3 Mar 2026 17:01 UTC
28 points
0 comments2 min readLW link

LASR Labs Sum­mer 2026 ap­pli­ca­tions are open!

3 Mar 2026 15:42 UTC
25 points
0 comments3 min readLW link

AI com­pa­nies and the 99% lethal au­tonomous weapons myth

User_Luke3 Mar 2026 12:18 UTC
−1 points
1 comment2 min readLW link

I’m con­fused by the change in the METR trend

Expertium3 Mar 2026 11:30 UTC
46 points
17 comments2 min readLW link

Zurich AI Safety is hiring a Director

3 Mar 2026 10:29 UTC
21 points
0 comments3 min readLW link

Game Rec­og­nizes Game

eva_3 Mar 2026 10:09 UTC
105 points
15 comments12 min readLW link

Mon­day AI Radar #15

Against Moloch3 Mar 2026 5:23 UTC
13 points
1 comment7 min readLW link

Me­mory De­cod­ing Jour­nal Club: En­gram cell con­nec­tivity as a mechanism for in­for­ma­tion en­cod­ing and mem­ory func­tion

Devin Ward3 Mar 2026 1:32 UTC
3 points
0 comments1 min readLW link

In-con­text learn­ing of rep­re­sen­ta­tions can be ex­plained by in­duc­tion circuits

Andy Arditi2 Mar 2026 23:58 UTC
51 points
0 comments9 min readLW link
(iclr-blogposts.github.io)

Sin­gle Direc­tion vs Low-Rank Re­fusal in Small LLMs

IvanC2 Mar 2026 23:14 UTC
12 points
0 comments8 min readLW link

Be­ing am­bi­tious in soulful altruism

pandamonium2 Mar 2026 21:15 UTC
5 points
0 comments4 min readLW link

Notes on the “Heart of Dark­ness”

dominicq2 Mar 2026 20:11 UTC
5 points
0 comments4 min readLW link
(sundaystopwatch.eu)

[Question] Can LLM chat be less pro­lix?

jbash2 Mar 2026 19:54 UTC
21 points
9 comments2 min readLW link

Ep­stein and my world model

Eye You2 Mar 2026 18:15 UTC
43 points
13 comments1 min readLW link

CLR Sum­mer Re­search Fel­low­ship 2026

Tristan Cook2 Mar 2026 18:03 UTC
32 points
0 comments7 min readLW link

War Claude

PeterMcCluskey2 Mar 2026 17:23 UTC
52 points
2 comments3 min readLW link
(bayesianinvestor.com)

Sec­re­tary of War Tweets That An­thropic is Now a Sup­ply Chain Risk

Zvi2 Mar 2026 13:20 UTC
81 points
15 comments91 min readLW link
(thezvi.wordpress.com)

“ball brain­teaser 4 color beads slide ru­bics cube” and mean­ing-making

flying buttress2 Mar 2026 11:47 UTC
16 points
1 comment3 min readLW link

Ex­plain­ing un­de­sir­able model be­hav­ior: (How) can in­fluence func­tions help?

2 Mar 2026 11:30 UTC
18 points
0 comments3 min readLW link

[Question] If ‘bad guys’ don’t pause, do you?

Remmelt2 Mar 2026 7:24 UTC
24 points
3 comments1 min readLW link

How to De­sign En­vi­ron­ments for Un­der­stand­ing Model Motives

2 Mar 2026 7:14 UTC
51 points
0 comments10 min readLW link

Con­text Aware­ness: Con­sti­tu­tional AI can miti­gate Emer­gent Misalignement

2 Mar 2026 5:21 UTC
25 points
18 comments36 min readLW link

Con­tro­versy sur­round­ing Molt­book ob­scures its very real, novel, un­ex­pressed and rapidly emerg­ing safety risks

Lloyd Rhodes-Brandon2 Mar 2026 2:05 UTC
13 points
1 comment5 min readLW link

An Open Let­ter to the Depart­ment of War and Congress

Gordon Seidoh Worley1 Mar 2026 18:31 UTC
62 points
5 comments1 min readLW link

An Em­piri­cal Re­view of the An­i­mal Harm Benchmark

lukasgebhard1 Mar 2026 18:20 UTC
16 points
0 comments1 min readLW link
(forum.effectivealtruism.org)

In­tro­duc­ing and Dep­re­cat­ing WoFBench

jefftk1 Mar 2026 18:20 UTC
78 points
1 comment3 min readLW link
(www.jefftk.com)