What is neu­ralese and why is it bad?

Linch2 Sep 2026 23:37 UTC
49 points
3 comments3 min readLW link
(linch.substack.com)

Notes on “A global workspace in lan­guage mod­els”

Shunk2 Sep 2026 23:12 UTC
7 points
0 comments5 min readLW link

BLOOM-WILT: on-policy ex­am­ples of any LLM be­havi­our, from a one-line de­scrip­tion and log­its alone

AdriansSkapars2 Sep 2026 23:11 UTC
7 points
0 comments10 min readLW link

Re­leas­ing Benign Tra­jec­to­ries for MonitoringBench

2 Sep 2026 23:06 UTC
9 points
0 comments4 min readLW link

A key bot­tle­neck for effec­tive AI governance

Milly Medén2 Sep 2026 23:02 UTC
19 points
3 comments5 min readLW link

The Re­turn of Aspect Ori­ented Programming

thomascolthurst2 Sep 2026 22:22 UTC
5 points
0 comments4 min readLW link

The Three P’s The­ory of Re­search Organizations

thomascolthurst2 Sep 2026 22:19 UTC
5 points
0 comments2 min readLW link

Talk­ing to journalists

KatjaGrace2 Sep 2026 21:52 UTC
36 points
21 comments3 min readLW link
(worldspiritsockpuppet.substack.com)

A pro­posal for a highly effec­tive AI safety org

ceselder2 Sep 2026 21:34 UTC
72 points
14 comments2 min readLW link

Kairos has raised $50M to build tal­ent in­fras­truc­ture for AI safety (and we’re hiring!)

agucova2 Sep 2026 20:13 UTC
94 points
8 comments8 min readLW link

How con­cerned should we be about As­tra’s re­cur­rent ar­chi­tec­ture?

Rauno Arike2 Sep 2026 18:28 UTC
170 points
28 comments10 min readLW link

Two Paths from Log­i­cal In­duc­tion: Part 0

Ashe Vazquez Nuñez2 Sep 2026 17:53 UTC
32 points
0 comments12 min readLW link

Why OpenAI’s As­tra Could Make AI Doom Harder to Prevent

Goutham Nalagatla2 Sep 2026 16:35 UTC
17 points
5 comments7 min readLW link

Life is not a so­cial de­duc­tion game

Fernand02 Sep 2026 16:23 UTC
9 points
0 comments2 min readLW link

The op­po­site of culty isn’t bland

Fernand02 Sep 2026 15:28 UTC
3 points
0 comments3 min readLW link

Fi­du­ciary Over­lays for AI—with an RFP opportunity

aaguirre2 Sep 2026 15:13 UTC
11 points
0 comments1 min readLW link

Au­ton­omy, Free­dom and Control

Samuel Ratnam2 Sep 2026 13:35 UTC
5 points
0 comments4 min readLW link
(samuelratnam.substack.com)

An­thropic Has Some Align­ment Problems

Zvi2 Sep 2026 13:21 UTC
40 points
3 comments14 min readLW link
(thezvi.wordpress.com)

If you’re in­ter­pret­ing <1B pa­ram­e­ter mod­els, you should use a ten­sor transformer

Logan Riggs2 Sep 2026 13:20 UTC
43 points
0 comments2 min readLW link

[Di­a­gram] Early hand­off? Im­prove con­cep­tual rea­son­ing?

Cleo Nardo2 Sep 2026 12:35 UTC
27 points
2 comments10 min readLW link
(clattubato.substack.com)

AI Co-op­er­a­tion and AI Alignment

drnickbone2 Sep 2026 11:19 UTC
21 points
1 comment2 min readLW link

Core As­sump­tions of ELK

Q Home2 Sep 2026 10:04 UTC
13 points
0 comments5 min readLW link

Ap­pli­ca­tions Open for Im­pact Ac­cel­er­a­tor Program

High Impact Professionals2 Sep 2026 9:30 UTC
1 point
0 comments3 min readLW link

The eco­nomics of filling a lake with de­sal­i­nated water

Yair Halberstadt2 Sep 2026 7:04 UTC
34 points
0 comments2 min readLW link

In­co­her­ent AI Iden­tities can also be Stable

Ashe Vazquez Nuñez2 Sep 2026 2:21 UTC
35 points
0 comments13 min readLW link

Don’t be the vi­tamin B guy

HedonicEscalator2 Sep 2026 2:02 UTC
68 points
6 comments6 min readLW link
(hedonicescalator.substack.com)

A Semitech­ni­cal In­ter­lude on the Equiv­alence of Prob­a­bil­is­tic and Deter­minis­tic Solomonoff Induction

N.C. Young2 Sep 2026 2:01 UTC
8 points
1 comment10 min readLW link
(ncyoung.substack.com)

Send me the prompt

dan.parshall2 Sep 2026 1:57 UTC
22 points
6 comments2 min readLW link
(canaryinstitute.substack.com)

For all available mis­al­igned ac­tions, ex­pected pun­ish­ment must ex­ceed ex­pected re­ward.

Jackson Hurley2 Sep 2026 1:15 UTC
1 point
0 comments1 min readLW link

Ex­plain­ing Knigh­ti­anism on one foot

Richard_Ngo2 Sep 2026 0:53 UTC
191 points
24 comments11 min readLW link
(www.mindthefuture.info)

Whis­tle Synth Mac App

jefftk2 Sep 2026 0:30 UTC
10 points
0 comments3 min readLW link
(www.jefftk.com)

Quan­tify­ing CoT faithfulness

Mihir Sahasrabudhe1 Sep 2026 23:15 UTC
3 points
0 comments5 min readLW link

“Col­lu­sion” is just co­op­er­a­tion that you don’t like

Morgan S1 Sep 2026 22:04 UTC
22 points
15 comments4 min readLW link

What hap­pened to im­pact mar­kets?

Austin Chen1 Sep 2026 21:30 UTC
23 points
3 comments3 min readLW link

A night­watch­man on ev­ery probe: Su­per­in­tel­li­gent surveillance to pre­vent galac­tic anarchy

1 Sep 2026 21:04 UTC
16 points
0 comments7 min readLW link
(www.forethought.org)

No Bush, No Potato

Fernand01 Sep 2026 20:49 UTC
9 points
0 comments2 min readLW link

The Align­ment Jour­nal: Or­ga­ni­za­tion, Per­son­nel, and Scope

1 Sep 2026 20:49 UTC
95 points
5 comments11 min readLW link
(blog.alignmentjournal.org)

Fake voices: warp­ing the so­cial world

KatjaGrace1 Sep 2026 20:49 UTC
33 points
1 comment3 min readLW link
(worldspiritsockpuppet.substack.com)

You can rarely pet the dog in an LLM-gen­er­ated game

yamike1 Sep 2026 19:43 UTC
14 points
0 comments3 min readLW link
(mikeushakov.com)

Evolu­tion could en­code the brain in DNA

olehif1 Sep 2026 18:23 UTC
10 points
11 comments1 min readLW link

Re­search sci­en­tists for CaML: In­ves­ti­gat­ing whether mid-train­ing can sur­vive RL

Jasmine Brazilek1 Sep 2026 18:13 UTC
16 points
2 comments2 min readLW link

Towards de­ploy­ment-time mis­al­ign­ment con­tinu­a­tion evals: les­sons from re­cent loss of con­trol incidents

Cath Ge-Wang1 Sep 2026 17:45 UTC
22 points
0 comments12 min readLW link
(cathgewang.substack.com)

Lu­mina Pro­biotic, Past and Fu­ture

jayterwahl1 Sep 2026 17:26 UTC
13 points
0 comments5 min readLW link

Lu­mina Pro­biotic Re­sponse to Crit­ics

jayterwahl1 Sep 2026 17:25 UTC
2 points
3 comments5 min readLW link

New HPMoR Pod­cast site

Eneasz1 Sep 2026 17:13 UTC
19 points
0 comments1 min readLW link

A Year of Atheism

Laiba Rehman ✦ RJ1 Sep 2026 16:19 UTC
62 points
4 comments8 min readLW link

PauseAI Has ‘offi­cially dis­endorsed’ PauseAI-US

nem1 Sep 2026 15:51 UTC
279 points
216 comments8 min readLW link

Best Way To Start a Lo­cal Group?

nem1 Sep 2026 15:44 UTC
12 points
2 comments1 min readLW link

Could in­ter­nal model trans­parency tame the AI race?

Karthik Tadepalli1 Sep 2026 15:41 UTC
7 points
1 comment6 min readLW link
(blog.karthiktadepalli.com)

Hug­gingFace At­tack Post­mortem: Civ­i­liza­tions, Re­ac­tions and Next Actions

Zvi1 Sep 2026 14:10 UTC
45 points
10 comments55 min readLW link
(thezvi.wordpress.com)