Ab­strac­tion Equivocation

WillPetillo3 Sep 2026 23:28 UTC
20 points
2 comments4 min readLW link

A Ther­a­pist for Peo­ple Who Think the World Might End: An In­ter­view with Daystar Eld (Da­mon Sasi)

JohnGreer3 Sep 2026 22:14 UTC
7 points
0 comments49 min readLW link
(youtu.be)

Biop­unk Green­house: Trees, Biochar, and Open Plant Biotechnology

Byron Lee3 Sep 2026 21:56 UTC
2 points
0 comments3 min readLW link

Work at Man­i­fund or Mox

Carol N3 Sep 2026 21:46 UTC
3 points
0 comments3 min readLW link
(manifund.substack.com)

Cat-Bel­ling Problems

Eliezer Yudkowsky3 Sep 2026 21:20 UTC
321 points
92 comments21 min readLW link

Cor­rigi­bil­ity Re­search Fund Gran­tees (Round 1)

Max Harms3 Sep 2026 20:58 UTC
27 points
0 comments8 min readLW link

How I’m Eval­u­at­ing Cor­rigi­bil­ity Grant Applications

Max Harms3 Sep 2026 20:58 UTC
46 points
0 comments9 min readLW link

The bias still mak­ing some ex­perts un­der­es­ti­mate LLMs

Steff3 Sep 2026 20:11 UTC
18 points
0 comments4 min readLW link

We need a global train­ing cut­off of April 2026

Ben Livengood3 Sep 2026 19:12 UTC
17 points
4 comments2 min readLW link

Fron­tier AI Shops Should Be Fil­ter­ing and Gen­er­at­ing AI Dis­course Dur­ing Pretraining

hillz3 Sep 2026 18:38 UTC
8 points
3 comments3 min readLW link

Steer­ing to­wards “au­to­mated grad­ing” de­grades alignment

3 Sep 2026 18:17 UTC
154 points
34 comments6 min readLW link

Cached cor­re­lated ran­dom­iza­tion: a tweak to UDT in ad­ver­sar­ial games

cousin_it3 Sep 2026 18:03 UTC
27 points
4 comments2 min readLW link

From safety re­search prompt to cross-model uni­ver­sal jailbreak

richbc3 Sep 2026 17:42 UTC
42 points
1 comment18 min readLW link

Sen. Bernie San­ders (I-VT) and Rep. Greg Casar (D-TX) in­tro­duce leg­is­la­tion to ban Ar­tifi­cial Su­per­in­tel­li­gence and tem­porar­ily pause ad­vanced AI de­vel­op­ment

Matrice Jacobine3 Sep 2026 14:41 UTC
225 points
31 comments2 min readLW link
(www.sanders.senate.gov)

AI #184: Post Post Mortem

Zvi3 Sep 2026 14:30 UTC
29 points
1 comment57 min readLW link
(thezvi.wordpress.com)

Im­prov­ing Au­dit Real­ism With In­fer­ence Time Com­pute and De­ploy­ment Scaffolds

3 Sep 2026 14:16 UTC
19 points
0 comments8 min readLW link

ACX Every­where Fall 26′ in Ho Chi Minh city

cygnus3 Sep 2026 12:18 UTC
2 points
0 comments1 min readLW link

Build­ing a High Ta­lent Den­sity Organization

Henry Papadatos3 Sep 2026 12:08 UTC
3 points
0 comments3 min readLW link
(henrypapadatos.com)

Per­sonal Effort to Re­duce Biorisk

jefftk3 Sep 2026 1:50 UTC
36 points
3 comments2 min readLW link
(www.jefftk.com)

Satis­fy­ing Cu­ri­os­ity Isn’t the Same as Understanding

CstineSublime3 Sep 2026 1:04 UTC
14 points
6 comments1 min readLW link

Re­s­olu­tion has a new Agent Foun­da­tions team

Jeremy Gillen3 Sep 2026 0:53 UTC
244 points
4 comments2 min readLW link
(resolution.org)

What is neu­ralese and why is it bad?

Linch2 Sep 2026 23:37 UTC
49 points
3 comments3 min readLW link
(linch.substack.com)

Notes on “A global workspace in lan­guage mod­els”

Shunk2 Sep 2026 23:12 UTC
7 points
0 comments5 min readLW link

BLOOM-WILT: on-policy ex­am­ples of any LLM be­havi­our, from a one-line de­scrip­tion and log­its alone

AdriansSkapars2 Sep 2026 23:11 UTC
7 points
0 comments10 min readLW link

Re­leas­ing Benign Tra­jec­to­ries for MonitoringBench

2 Sep 2026 23:06 UTC
9 points
0 comments4 min readLW link

A key bot­tle­neck for effec­tive AI governance

Milly Medén2 Sep 2026 23:02 UTC
19 points
3 comments5 min readLW link

The Re­turn of Aspect Ori­ented Programming

thomascolthurst2 Sep 2026 22:22 UTC
5 points
0 comments4 min readLW link

The Three P’s The­ory of Re­search Organizations

thomascolthurst2 Sep 2026 22:19 UTC
5 points
0 comments2 min readLW link

Talk­ing to journalists

KatjaGrace2 Sep 2026 21:52 UTC
36 points
21 comments3 min readLW link
(worldspiritsockpuppet.substack.com)

A pro­posal for a highly effec­tive AI safety org

ceselder2 Sep 2026 21:34 UTC
72 points
14 comments2 min readLW link

Kairos has raised $50M to build tal­ent in­fras­truc­ture for AI safety (and we’re hiring!)

agucova2 Sep 2026 20:13 UTC
94 points
8 comments8 min readLW link

How con­cerned should we be about As­tra’s re­cur­rent ar­chi­tec­ture?

Rauno Arike2 Sep 2026 18:28 UTC
170 points
28 comments10 min readLW link

Two Paths from Log­i­cal In­duc­tion: Part 0

Ashe Vazquez Nuñez2 Sep 2026 17:53 UTC
32 points
0 comments12 min readLW link

Why OpenAI’s As­tra Could Make AI Doom Harder to Prevent

Goutham Nalagatla2 Sep 2026 16:35 UTC
17 points
5 comments7 min readLW link

Life is not a so­cial de­duc­tion game

Fernand02 Sep 2026 16:23 UTC
9 points
0 comments2 min readLW link

The op­po­site of culty isn’t bland

Fernand02 Sep 2026 15:28 UTC
3 points
0 comments3 min readLW link

Fi­du­ciary Over­lays for AI—with an RFP opportunity

aaguirre2 Sep 2026 15:13 UTC
11 points
0 comments1 min readLW link

Au­ton­omy, Free­dom and Control

Samuel Ratnam2 Sep 2026 13:35 UTC
5 points
0 comments4 min readLW link
(samuelratnam.substack.com)

An­thropic Has Some Align­ment Problems

Zvi2 Sep 2026 13:21 UTC
40 points
3 comments14 min readLW link
(thezvi.wordpress.com)

If you’re in­ter­pret­ing <1B pa­ram­e­ter mod­els, you should use a ten­sor transformer

Logan Riggs2 Sep 2026 13:20 UTC
43 points
0 comments2 min readLW link

[Di­a­gram] Early hand­off? Im­prove con­cep­tual rea­son­ing?

Cleo Nardo2 Sep 2026 12:35 UTC
27 points
2 comments10 min readLW link
(clattubato.substack.com)

AI Co-op­er­a­tion and AI Alignment

drnickbone2 Sep 2026 11:19 UTC
21 points
1 comment2 min readLW link

Core As­sump­tions of ELK

Q Home2 Sep 2026 10:04 UTC
13 points
0 comments5 min readLW link

Ap­pli­ca­tions Open for Im­pact Ac­cel­er­a­tor Program

High Impact Professionals2 Sep 2026 9:30 UTC
1 point
0 comments3 min readLW link

The eco­nomics of filling a lake with de­sal­i­nated water

Yair Halberstadt2 Sep 2026 7:04 UTC
34 points
0 comments2 min readLW link

In­co­her­ent AI Iden­tities can also be Stable

Ashe Vazquez Nuñez2 Sep 2026 2:21 UTC
35 points
0 comments13 min readLW link

Don’t be the vi­tamin B guy

HedonicEscalator2 Sep 2026 2:02 UTC
68 points
6 comments6 min readLW link
(hedonicescalator.substack.com)

A Semitech­ni­cal In­ter­lude on the Equiv­alence of Prob­a­bil­is­tic and Deter­minis­tic Solomonoff Induction

N.C. Young2 Sep 2026 2:01 UTC
8 points
1 comment10 min readLW link
(ncyoung.substack.com)

Send me the prompt

dan.parshall2 Sep 2026 1:57 UTC
22 points
6 comments2 min readLW link
(canaryinstitute.substack.com)

For all available mis­al­igned ac­tions, ex­pected pun­ish­ment must ex­ceed ex­pected re­ward.

Jackson Hurley2 Sep 2026 1:15 UTC
1 point
0 comments1 min readLW link