Videogames for Rationalists

Adam Newgas8 Mar 2026 21:59 UTC
26 points
9 comments1 min readLW link

Fake Up­dates

Algon8 Mar 2026 21:14 UTC
17 points
3 comments2 min readLW link
(algon33.substack.com)

Re­cre­ation of EA-Pioneer Igor Kiriluk

avturchin8 Mar 2026 18:51 UTC
56 points
4 comments2 min readLW link

Don’t ac­cuse your in­ter­locu­tor of mak­ing ar­gu­ments that aren’t rooted in evidence

TFD8 Mar 2026 17:39 UTC
−1 points
2 comments2 min readLW link
(www.thefloatingdroid.com)

1999 JavaScript and 2025 AI: Same Cir­cus, Differ­ent Tent

Scott S Nelson8 Mar 2026 15:40 UTC
−8 points
0 comments4 min readLW link

How to Get Kids In­ter­ested in Science and Scien­tific Reasoning

Rami Rustom8 Mar 2026 14:41 UTC
−1 points
0 comments2 min readLW link

On The In­de­pen­dence Axiom

Ihor Kendiukhov8 Mar 2026 14:38 UTC
336 points
93 comments23 min readLW link

Pri­vacy, Hon­esty, Im­perfect Glo­ma­riz­ing: Pick two

shelvacu8 Mar 2026 14:26 UTC
6 points
7 comments1 min readLW link

So­lar Storms

Croissanthology8 Mar 2026 14:04 UTC
173 points
43 comments12 min readLW link

Draft Moskovitz: The Best Last Hope for Con­struc­tive AI Governance

Oliver Kuperman8 Mar 2026 13:42 UTC
6 points
0 comments11 min readLW link

The Law of Pos­i­tive-Sum Badness

Davidmanheim8 Mar 2026 13:27 UTC
51 points
0 comments10 min readLW link

Does re­search from mat­spro­gram.org/​re­search aim to help re­duce P(doom)? Let’s find out! (with Gem­ini 3.1 Pro) Part 1

Zabor8 Mar 2026 13:20 UTC
1 point
0 comments7 min readLW link

Open let­ter to doomers

delphix8 Mar 2026 8:12 UTC
−14 points
10 comments4 min readLW link

Co­op­er­a­tion Without Kind­ness or Strategy

seank8 Mar 2026 3:36 UTC
3 points
2 comments3 min readLW link
(forum.effectivealtruism.org)

Why Many Am­bi­tious (and Altru­is­tic) Peo­ple Prob­a­bly Un­der­value Their Happiness

emily.fan8 Mar 2026 2:31 UTC
2 points
0 comments7 min readLW link

The cur­rent SOTA model was re­leased with­out safety evals

8 Mar 2026 1:51 UTC
111 points
12 comments5 min readLW link

Pro­posal For Cryp­to­graphic Method to Ri­gor­ously Ver­ify LLM Prompt Experiments

weberr137 Mar 2026 21:09 UTC
5 points
0 comments2 min readLW link

The first con­firmed in­stance of an LLM go­ing rogue for in­stru­men­tal rea­sons in a real-world set­ting has oc­curred, buried in an Alibaba pa­per about a new train­ing pipeline.

lilkim20257 Mar 2026 20:18 UTC
71 points
22 comments2 min readLW link

[Question] When has fore­cast­ing been use­ful for you?

sanyer7 Mar 2026 19:50 UTC
14 points
4 comments1 min readLW link

Can gov­ern­ments quickly and cheaply slow AI train­ing?

joshc7 Mar 2026 19:11 UTC
64 points
9 comments14 min readLW link

Did I Catch Claude Cheat­ing?

weberr137 Mar 2026 6:08 UTC
14 points
2 comments4 min readLW link

D&D.Sci Re­lease Day: Top­ple the Tower!

aphyer7 Mar 2026 2:48 UTC
29 points
17 comments2 min readLW link

AI Safety Needs Startups

7 Mar 2026 1:27 UTC
10 points
4 comments12 min readLW link
(blog.bluedot.org)

CHAI 2026 Work­shop: Open Call for Posters!

Sarah Otis7 Mar 2026 1:17 UTC
2 points
0 comments1 min readLW link
(workshop.humancompatible.ai)

More is differ­ent for intelligence

7 Mar 2026 0:02 UTC
17 points
0 comments2 min readLW link
(fulcruminc.substack.com)

Your Causal Vari­ables Are Irre­ducibly Subjective

David Reber6 Mar 2026 22:59 UTC
45 points
3 comments5 min readLW link

Mox is the largest AI Safety com­mu­nity space in San Fran­cisco. We’re fundrais­ing!

Rachel Shu6 Mar 2026 22:07 UTC
30 points
0 comments8 min readLW link

Self-At­tri­bu­tion Bias: When AI Mon­i­tors Go Easy on Themselves

6 Mar 2026 21:54 UTC
44 points
6 comments6 min readLW link

Thoughts on the Pause AI protest

philh6 Mar 2026 21:50 UTC
136 points
18 comments7 min readLW link
(reasonableapproximation.net)

Pod­cast: Jeremy Howard is bear­ish on LLMs

Steven Byrnes6 Mar 2026 21:39 UTC
89 points
24 comments5 min readLW link
(www.youtube.com)

La­tent Rea­son­ing Sprint #1: Tuned Lens and Logit Lens on CODI

Realmbird6 Mar 2026 18:36 UTC
7 points
1 comment4 min readLW link

An­thropic Offi­cially, Ar­bi­trar­ily and Capri­ciously Des­ig­nated a Sup­ply Chain Risk

Zvi6 Mar 2026 18:10 UTC
68 points
1 comment19 min readLW link
(thezvi.wordpress.com)

The Elect

Tomás B.6 Mar 2026 15:34 UTC
100 points
1 comment16 min readLW link
(open.substack.com)

Play­ing Pos­sum: The Vari­abil­ity Hypothesis

rba6 Mar 2026 14:48 UTC
22 points
1 comment6 min readLW link
(goflaw.substack.com)

Shap­ing the ex­plo­ra­tion of the mo­ti­va­tion-space mat­ters for AI safety

6 Mar 2026 14:43 UTC
85 points
16 comments10 min readLW link

A Com­po­si­tional Philos­o­phy of Science for Agent Foundations

Jonas Hallgren6 Mar 2026 8:40 UTC
30 points
1 comment13 min readLW link
(equilibria1.substack.com)

It Is Always Worth­while to En­gage in a Dis­cus­sion on the Public Web

Bowl of Cereal6 Mar 2026 6:21 UTC
−6 points
0 comments2 min readLW link

How I Han­dle Au­to­mated Programming

HunterJay6 Mar 2026 4:27 UTC
14 points
0 comments8 min readLW link

Rea­son­ing Models Strug­gle to Con­trol Their Chains of Thought

5 Mar 2026 22:37 UTC
76 points
9 comments3 min readLW link

Per­son­al­ity Self-Replicators

eggsyntax5 Mar 2026 20:30 UTC
173 points
51 comments10 min readLW link

Salient Direc­tions in AI Control

Bruce W. Lee5 Mar 2026 19:38 UTC
13 points
0 comments14 min readLW link
(brucewlee.com)

Models have lin­ear rep­re­sen­ta­tions of what tasks they like

OscarGilg5 Mar 2026 18:44 UTC
55 points
16 comments11 min readLW link

AI Safety Has 12 Months Left

mhdempsey5 Mar 2026 16:37 UTC
40 points
10 comments6 min readLW link
(mhdempsey.substack.com)

Have Amer­i­cans Be­come Less Violent Since 1980?

Benquo5 Mar 2026 16:11 UTC
78 points
6 comments14 min readLW link
(benjaminrosshoffman.com)

AI #158: The Depart­ment of War

Zvi5 Mar 2026 16:10 UTC
77 points
2 comments51 min readLW link
(thezvi.wordpress.com)

In­ves­ti­gat­ing Self-Fulfilling Misal­ign­ment and Col­lu­sion in AI Control

5 Mar 2026 15:05 UTC
15 points
0 comments5 min readLW link

Com­pu­ta­tion, Chess, and Lan­guage in Ar­tifi­cial In­tel­li­gence

Bill Benzon5 Mar 2026 12:57 UTC
6 points
0 comments3 min readLW link

Vibe Cod­ing crip­ples the mind

spookyuser5 Mar 2026 10:29 UTC
−12 points
4 comments4 min readLW link

Ra­tional Chess

8495 Mar 2026 9:57 UTC
5 points
19 comments2 min readLW link

A Be­havi­oural and Rep­re­sen­ta­tional Eval­u­a­tion of Goal-di­rect­ed­ness in Lan­guage Model Agents

5 Mar 2026 1:08 UTC
20 points
0 comments7 min readLW link