The Loss of Singularity

Benoit Henry16 Sep 2026 22:59 UTC
3 points
0 comments5 min readLW link

Re­duc­ing the Re­source Gap Between Lab and Ex­ter­nal Safety Researchers

16 Sep 2026 22:51 UTC
49 points
1 comment7 min readLW link

Thought an­chors don’t trans­fer be­tween models

therootof316 Sep 2026 22:49 UTC
8 points
0 comments9 min readLW link

Why is AI so un­reg­u­lated?

KatjaGrace16 Sep 2026 22:39 UTC
24 points
8 comments1 min readLW link
(worldspiritsockpuppet.substack.com)

The Most Im­por­tant City in AI Safety May Be Singapore

Ran Sun16 Sep 2026 22:35 UTC
11 points
0 comments3 min readLW link

Launch­ing RESI, the In­sti­tute for Re­spon­si­ble Superintelligence

adamology16 Sep 2026 22:18 UTC
10 points
1 comment3 min readLW link

What the Hug­ging Face In­ci­dent tells us about multi-agent interactions

Himnish116 Sep 2026 22:14 UTC
9 points
0 comments6 min readLW link

Always ask your agent to be honest

16 Sep 2026 22:01 UTC
11 points
0 comments13 min readLW link

Ob­sta­cles to the scal­able over­sight of auto-al­ign­ment research

16 Sep 2026 21:52 UTC
44 points
0 comments9 min readLW link

Find­ing het­ero­ge­neous agent swarms in the wild

Michael Flood16 Sep 2026 21:50 UTC
14 points
0 comments2 min readLW link
(mflood.substack.com)

If Any­one Builds It, Every­one Dies: One Year Closer

16 Sep 2026 20:22 UTC
338 points
18 comments8 min readLW link

Model or­ganisms (some­times) con­fess their mis­al­ign­ment when offered a deal

16 Sep 2026 18:38 UTC
40 points
5 comments17 min readLW link

55% of the US pub­lic is now aware of AI xrisk

otto.barten16 Sep 2026 17:31 UTC
28 points
0 comments1 min readLW link

How to Open Them Up – Part I

ValueShift Research16 Sep 2026 16:40 UTC
8 points
0 comments8 min readLW link

Microsoft AI’s “Hu­man­ist” CoC

Stephen Martin16 Sep 2026 16:10 UTC
30 points
7 comments4 min readLW link

Run­ning a Ba­sic Bal­lot Meetup

Screwtape16 Sep 2026 15:43 UTC
13 points
0 comments2 min readLW link

How to Run a Bal­lot Meetup

Taymon Beal16 Sep 2026 15:26 UTC
7 points
0 comments13 min readLW link

Trump Goes Full Hoax on AI Ex­is­ten­tial Risk

Zvi16 Sep 2026 15:10 UTC
56 points
14 comments26 min readLW link
(thezvi.wordpress.com)

J-space au­dit­ing might be unreliable

Kartikay Luthra16 Sep 2026 14:06 UTC
4 points
1 comment13 min readLW link

Should our jour­nal pub­lish AI-drafted manuscripts?

Dan MacKinlay16 Sep 2026 8:09 UTC
16 points
4 comments27 min readLW link
(danmackinlay.name)

Who com­putes?

Dan MacKinlay16 Sep 2026 8:08 UTC
11 points
3 comments18 min readLW link
(danmackinlay.name)

Phan­tom trans­fer works via ex­tremely sub­tle se­man­tic cues

16 Sep 2026 4:52 UTC
55 points
2 comments20 min readLW link

An Alien Mind

papetoast16 Sep 2026 2:52 UTC
2 points
0 comments1 min readLW link
(openai.com)

Re­search ac­cel­er­a­tion: The view in­side OpenAI

papetoast16 Sep 2026 2:52 UTC
4 points
0 comments1 min readLW link
(openai.com)

The Hug­ging Face in­ci­dent and the road ahead

papetoast16 Sep 2026 2:52 UTC
4 points
0 comments1 min readLW link
(openai.com)

Is METR A Mean­ingful Check On An­thropic?

SE Gyges16 Sep 2026 0:34 UTC
113 points
44 comments6 min readLW link
(www.verysane.ai)

Why would AI cause hu­man ex­tinc­tion?

jacob15 Sep 2026 23:49 UTC
7 points
0 comments10 min readLW link

Inoc­u­la­tion Mid­train­ing with Learned Neologisms

15 Sep 2026 21:52 UTC
76 points
3 comments6 min readLW link
(arxiv.org)

Quick notes from teach­ing tech­ni­cal pro­files how to talk in public

Camille B. 15 Sep 2026 21:52 UTC
210 points
7 comments6 min readLW link

Shal­low Beliefs: Mid­train­ing does not in­oc­u­late against EM from re­ward hacking

15 Sep 2026 21:50 UTC
46 points
3 comments3 min readLW link

We Should As­sume We Have One Chance At AI Legislation

Jamie Joyce15 Sep 2026 19:42 UTC
33 points
1 comment5 min readLW link

Why I’m do­ing the Su­san Calvin Project

Haoxing Du15 Sep 2026 19:06 UTC
20 points
0 comments4 min readLW link
(susancalvinproject.substack.com)

Any AI pause will have defec­tors. How to en­sure their in­car­cer­a­tion ac­tu­ally pre­vents them from covertly con­tribut­ing to AI re­search from be­hind bars?

Connor Williams15 Sep 2026 18:31 UTC
3 points
0 comments1 min readLW link
(connorsscratchpad.substack.com)

You Don’t Have to Trust the AI Labs (in or­der to take their call for reg­u­la­tion se­ri­ously)

ixotope15 Sep 2026 18:04 UTC
11 points
0 comments11 min readLW link

Co­op­er­a­tion with AIs seems to be a low-hang­ing fruit for bet­ter eval practices

Clément Dumas15 Sep 2026 17:48 UTC
136 points
27 comments8 min readLW link

Align­ment & Suc­ces­sion: Mo­ral­ity Lives in the Hu­man Individual

L Rudolf L15 Sep 2026 15:38 UTC
29 points
3 comments16 min readLW link

Why Fo­cus on Ex­tinc­tion?

alkjash15 Sep 2026 14:37 UTC
8 points
2 comments1 min readLW link
(radimentary.wordpress.com)

The Bad Guy With An AI Named Claude

Zvi15 Sep 2026 14:10 UTC
38 points
7 comments19 min readLW link
(thezvi.wordpress.com)

Moloch Does My Hair

spookyuser15 Sep 2026 10:31 UTC
19 points
0 comments19 min readLW link

For most peo­ple “in­tel­li­gence” is not goal achievement

Pato15 Sep 2026 7:52 UTC
13 points
2 comments1 min readLW link

How we might ac­tu­ally pace the fron­tier: A pro­posal for AI com­pa­nies to do pub­lic pac­ing ex­er­cises.

Dewi Erwan15 Sep 2026 5:36 UTC
9 points
0 comments4 min readLW link
(blog.dewierwan.com)

As­tra ap­pears to perform be­lief-prop­a­ga­tion-like in­fer­ence with­out CoT

MBaert15 Sep 2026 2:48 UTC
94 points
0 comments10 min readLW link

One co­or­di­nate breaks abliter­a­tion on Gemma-3

Abhishek Mishra15 Sep 2026 2:36 UTC
8 points
0 comments14 min readLW link

How to think about LLM effort

Tao Lin15 Sep 2026 2:31 UTC
17 points
0 comments2 min readLW link

Study 3: Steer­ing welfare-rele­vant di­rec­tions moved the rep­re­sen­ta­tion, but not [de­tectably] the behavior

ashesfall15 Sep 2026 2:11 UTC
7 points
1 comment17 min readLW link

Model Weight Exfil­tra­tion Seems Overrated

Vaniver15 Sep 2026 0:01 UTC
101 points
21 comments3 min readLW link

Im­prov­ing Psy­chi­a­tric Medicine Devel­op­ment with AI

CMLKevin14 Sep 2026 23:21 UTC
5 points
0 comments3 min readLW link

Align­ment Prob­lem Redux

Mateusz Bagiński14 Sep 2026 21:53 UTC
14 points
1 comment2 min readLW link

Self Inoculation

epicurus14 Sep 2026 20:45 UTC
30 points
4 comments12 min readLW link

Univer­sity of Cam­bridge – Col­lege EA Mee­tups Every­where Fall 2026

Screwtape14 Sep 2026 19:20 UTC
8 points
0 comments1 min readLW link