Pro­posal For Cryp­to­graphic Method to Ri­gor­ously Ver­ify LLM Prompt Experiments

weberr137 Mar 2026 21:09 UTC
5 points
0 comments2 min readLW link

The first con­firmed in­stance of an LLM go­ing rogue for in­stru­men­tal rea­sons in a real-world set­ting has oc­curred, buried in an Alibaba pa­per about a new train­ing pipeline.

lilkim20257 Mar 2026 20:18 UTC
71 points
22 comments2 min readLW link

[Question] When has fore­cast­ing been use­ful for you?

sanyer7 Mar 2026 19:50 UTC
14 points
4 comments1 min readLW link

Can gov­ern­ments quickly and cheaply slow AI train­ing?

joshc7 Mar 2026 19:11 UTC
64 points
9 comments14 min readLW link

Did I Catch Claude Cheat­ing?

weberr137 Mar 2026 6:08 UTC
14 points
2 comments4 min readLW link

D&D.Sci Re­lease Day: Top­ple the Tower!

aphyer7 Mar 2026 2:48 UTC
29 points
17 comments2 min readLW link

AI Safety Needs Startups

7 Mar 2026 1:27 UTC
10 points
4 comments12 min readLW link
(blog.bluedot.org)

CHAI 2026 Work­shop: Open Call for Posters!

Sarah Otis7 Mar 2026 1:17 UTC
2 points
0 comments1 min readLW link
(workshop.humancompatible.ai)

More is differ­ent for intelligence

7 Mar 2026 0:02 UTC
17 points
0 comments2 min readLW link
(fulcruminc.substack.com)

Your Causal Vari­ables Are Irre­ducibly Subjective

David Reber6 Mar 2026 22:59 UTC
45 points
3 comments5 min readLW link

Mox is the largest AI Safety com­mu­nity space in San Fran­cisco. We’re fundrais­ing!

Rachel Shu6 Mar 2026 22:07 UTC
30 points
0 comments8 min readLW link

Self-At­tri­bu­tion Bias: When AI Mon­i­tors Go Easy on Themselves

6 Mar 2026 21:54 UTC
44 points
6 comments6 min readLW link

Thoughts on the Pause AI protest

philh6 Mar 2026 21:50 UTC
136 points
18 comments7 min readLW link
(reasonableapproximation.net)

Pod­cast: Jeremy Howard is bear­ish on LLMs

Steven Byrnes6 Mar 2026 21:39 UTC
89 points
24 comments5 min readLW link
(www.youtube.com)

La­tent Rea­son­ing Sprint #1: Tuned Lens and Logit Lens on CODI

Realmbird6 Mar 2026 18:36 UTC
7 points
1 comment4 min readLW link

An­thropic Offi­cially, Ar­bi­trar­ily and Capri­ciously Des­ig­nated a Sup­ply Chain Risk

Zvi6 Mar 2026 18:10 UTC
68 points
1 comment19 min readLW link
(thezvi.wordpress.com)

The Elect

Tomás B.6 Mar 2026 15:34 UTC
100 points
1 comment16 min readLW link
(open.substack.com)

Play­ing Pos­sum: The Vari­abil­ity Hypothesis

rba6 Mar 2026 14:48 UTC
22 points
1 comment6 min readLW link
(goflaw.substack.com)

Shap­ing the ex­plo­ra­tion of the mo­ti­va­tion-space mat­ters for AI safety

6 Mar 2026 14:43 UTC
85 points
16 comments10 min readLW link

A Com­po­si­tional Philos­o­phy of Science for Agent Foundations

Jonas Hallgren6 Mar 2026 8:40 UTC
30 points
1 comment13 min readLW link
(equilibria1.substack.com)

It Is Always Worth­while to En­gage in a Dis­cus­sion on the Public Web

Bowl of Cereal6 Mar 2026 6:21 UTC
−6 points
0 comments2 min readLW link

How I Han­dle Au­to­mated Programming

HunterJay6 Mar 2026 4:27 UTC
14 points
0 comments8 min readLW link

Rea­son­ing Models Strug­gle to Con­trol Their Chains of Thought

5 Mar 2026 22:37 UTC
76 points
9 comments3 min readLW link

Per­son­al­ity Self-Replicators

eggsyntax5 Mar 2026 20:30 UTC
173 points
51 comments10 min readLW link

Salient Direc­tions in AI Control

Bruce W. Lee5 Mar 2026 19:38 UTC
13 points
0 comments14 min readLW link
(brucewlee.com)

Models have lin­ear rep­re­sen­ta­tions of what tasks they like

OscarGilg5 Mar 2026 18:44 UTC
55 points
16 comments11 min readLW link

AI Safety Has 12 Months Left

mhdempsey5 Mar 2026 16:37 UTC
40 points
10 comments6 min readLW link
(mhdempsey.substack.com)

Have Amer­i­cans Be­come Less Violent Since 1980?

Benquo5 Mar 2026 16:11 UTC
78 points
6 comments14 min readLW link
(benjaminrosshoffman.com)

AI #158: The Depart­ment of War

Zvi5 Mar 2026 16:10 UTC
77 points
2 comments51 min readLW link
(thezvi.wordpress.com)

In­ves­ti­gat­ing Self-Fulfilling Misal­ign­ment and Col­lu­sion in AI Control

5 Mar 2026 15:05 UTC
15 points
0 comments5 min readLW link

Com­pu­ta­tion, Chess, and Lan­guage in Ar­tifi­cial In­tel­li­gence

Bill Benzon5 Mar 2026 12:57 UTC
6 points
0 comments3 min readLW link

Vibe Cod­ing crip­ples the mind

spookyuser5 Mar 2026 10:29 UTC
−12 points
4 comments4 min readLW link

Ra­tional Chess

8495 Mar 2026 9:57 UTC
5 points
19 comments2 min readLW link

A Be­havi­oural and Rep­re­sen­ta­tional Eval­u­a­tion of Goal-di­rect­ed­ness in Lan­guage Model Agents

5 Mar 2026 1:08 UTC
20 points
0 comments7 min readLW link

Fe­bru­ary 2026 Links

nomagicpill5 Mar 2026 1:03 UTC
6 points
2 comments7 min readLW link
(nomagicpill.substack.com)

Text Com­pres­sion Can Help Se­cure Model Weights

Roy Rinberg4 Mar 2026 23:30 UTC
45 points
12 comments10 min readLW link

A sum­mary of Con­den­sa­tion and its re­la­tion to Nat­u­ral Latents

4 Mar 2026 22:22 UTC
86 points
0 comments10 min readLW link

Maybe there’s a pat­tern here?

dynomight4 Mar 2026 20:32 UTC
193 points
45 comments7 min readLW link

Gem­ini 3.1 Pro Aces Bench­marks, I Suppose

Zvi4 Mar 2026 20:10 UTC
27 points
2 comments9 min readLW link
(thezvi.wordpress.com)

Gen Z and AI: Ed­u­ca­tion, Well-be­ing, and the Labour Mar­ket

Jason Hung4 Mar 2026 19:19 UTC
9 points
0 comments11 min readLW link

Is GDP a Kind of Fac­tory?

Benquo4 Mar 2026 18:00 UTC
59 points
3 comments10 min readLW link
(benjaminrosshoffman.com)

Make Pow­er­ful Machines Verifiable

Naci Cankaya4 Mar 2026 14:20 UTC
22 points
4 comments4 min readLW link

How a Pinky Promise once stopped a war in the Mid­dle East.

positivesum4 Mar 2026 13:30 UTC
29 points
1 comment4 min readLW link

Split Per­son­al­ity Train­ing can de­tect Align­ment Faking

Florian_Dietz4 Mar 2026 11:49 UTC
39 points
0 comments6 min readLW link

Physics of RL: Toy scal­ing laws for the emer­gence of re­ward-seeking

Alex Meinke4 Mar 2026 8:12 UTC
120 points
9 comments10 min readLW link

Sa­cred val­ues of fu­ture AIs

Cleo Nardo4 Mar 2026 7:47 UTC
58 points
4 comments5 min readLW link

OpenAI’s surveillance lan­guage has many po­ten­tial loop­holes and they can do better

Tom Smith4 Mar 2026 4:25 UTC
153 points
4 comments10 min readLW link

Mass surveillance, red lines, and a crazy weekend

Boaz Barak4 Mar 2026 4:24 UTC
35 points
26 comments5 min readLW link

Lie To Me, But At Least Don’t Bullshit

Czynski4 Mar 2026 2:20 UTC
20 points
4 comments5 min readLW link
(dangeroussincerity.substack.com)

[Question] LLM co­her­en­ti­za­tion as an ob­vi­ous low-hang­ing fruit to try?

Épiphanie Gédéon4 Mar 2026 0:59 UTC
26 points
2 comments2 min readLW link