A nor­mal Fri­day in 2042

RobinHa11 Sep 2026 22:29 UTC
21 points
0 comments13 min readLW link

My recom­mended re­sources for AI safety, al­ign­ment, and ex­is­ten­tial risks

Lysandre Terrisse11 Sep 2026 22:19 UTC
11 points
0 comments2 min readLW link

Ap­pendix: Re­pro­duc­tion of the OpenAI-Hug­gingFace Incident

Stewart Slocum11 Sep 2026 21:57 UTC
15 points
0 comments13 min readLW link

Con­sider pos­i­tive feed­back loops

Pato11 Sep 2026 21:55 UTC
7 points
0 comments3 min readLW link

Align­ment Hierarchy

Lucina11 Sep 2026 21:26 UTC
2 points
0 comments4 min readLW link

What Hap­pens Now? Fore­cast­ing the Fal­lout from the Hug­ging Face Incident

ChristianWilliams11 Sep 2026 19:22 UTC
10 points
0 comments10 min readLW link
(metaculus.substack.com)

As­tra’s no-CoT limits track spec­u­la­tive depth, not step count

MBaert11 Sep 2026 18:05 UTC
101 points
4 comments10 min readLW link

Or­ga­niz­ing a Fermi mod­el­ing ‘hack’ on the cost of cul­tured meat; hiring a co-coordinator

david reinstein11 Sep 2026 18:01 UTC
6 points
0 comments1 min readLW link

Caroline Elli­son has joined Manifund

11 Sep 2026 17:44 UTC
9 points
8 comments3 min readLW link

CoT con­trol­la­bil­ity evals seem very un­der-elicited

Jozdien11 Sep 2026 17:12 UTC
63 points
3 comments4 min readLW link

Lo­cal Fac­tor Graph Debate

Alexander Heckett11 Sep 2026 16:59 UTC
24 points
0 comments9 min readLW link

Post-AGI, we are all jobless aristocrats

djbinder11 Sep 2026 16:53 UTC
34 points
9 comments3 min readLW link
(defensesindepth.bio)

AIRO: Au­to­mated fore­casts of catas­trophic risks

Nick Merrill11 Sep 2026 16:29 UTC
12 points
0 comments3 min readLW link

SFT Also Drives Safety Eval Re­sults in Olmo 3

Finn Cairns11 Sep 2026 16:20 UTC
27 points
0 comments1 min readLW link
(secondlookresearch.com)

The Golden Age of Impact

Bentham's Bulldog11 Sep 2026 15:47 UTC
−2 points
3 comments3 min readLW link

Con­trolAI’s Creator Outreach

11 Sep 2026 14:54 UTC
26 points
0 comments10 min readLW link
(blog.controlai.org)

Ja­cob Coxon Warns of Hu­man Ex­tinc­tion and Trig­gers a Prefer­ence Cascade

Zvi11 Sep 2026 14:40 UTC
85 points
2 comments43 min readLW link
(thezvi.wordpress.com)

Why Hug­gingFace Hasn’t Shifted my P(Doom)

Josh Snider11 Sep 2026 13:54 UTC
9 points
0 comments3 min readLW link

The Ex­tinc­tion Risk Prefer­ence Cas­cade: Quotes

Zvi11 Sep 2026 13:50 UTC
27 points
0 comments16 min readLW link
(thezvi.wordpress.com)

Gen­er­al­ized UDT 1.0 tiling

Roman Malov11 Sep 2026 12:28 UTC
18 points
0 comments6 min readLW link

Re­place Net Me­ter­ing With Batteries

jefftk11 Sep 2026 11:50 UTC
14 points
6 comments6 min readLW link
(www.jefftk.com)

Could Re­s­olu­tion Help Build Align­ment Re­search as a Dis­ci­pline?

IanWS11 Sep 2026 8:49 UTC
11 points
0 comments3 min readLW link

OpenAI-Hug­gingFace: A Re­pro­duc­tion & Les­sons for Align­ment Testing

Stewart Slocum11 Sep 2026 8:48 UTC
77 points
2 comments12 min readLW link

We need good evals for ac­ti­va­tion faithfulness

11 Sep 2026 8:20 UTC
25 points
0 comments2 min readLW link

Ques­tions for the “New En­light­en­ment” in the Age of AGI

Jordan Arel11 Sep 2026 5:49 UTC
8 points
0 comments4 min readLW link

(Ques­tion) When should I cross-post on LessWrong vs. EA Fo­rum only?

Jordan Arel11 Sep 2026 5:37 UTC
4 points
1 comment1 min readLW link

Is func­tional welfare speak­able?

11 Sep 2026 4:59 UTC
6 points
0 comments1 min readLW link
(latentminds.org)

Vol­un­tary Grad­ual Disem­pow­er­ment in the Ju­di­ciary/​Le­gal Sys­tem

Caleb Horn11 Sep 2026 4:00 UTC
14 points
0 comments1 min readLW link

Which char­ac­ter are we eval­u­at­ing? Per­sona sta­bil­ity and AI welfare

Joshua Fonseca Rivera11 Sep 2026 3:04 UTC
15 points
0 comments5 min readLW link

I won­der how it ends...

Lucina11 Sep 2026 2:00 UTC
6 points
0 comments1 min readLW link

Per­spec­tives in favour of im­prov­ing con­cep­tual rea­son­ing capabilities

Chi Nguyen11 Sep 2026 1:32 UTC
17 points
2 comments7 min readLW link

Okay, fine. I’ll try Substack

Alex_Altair11 Sep 2026 1:13 UTC
19 points
0 comments3 min readLW link
(alexaltair.substack.com)

How a cold email got the Fin­nish gov­ern­ment to re­spond on su­per­in­tel­li­gence regulation

Josh Thorsteinson11 Sep 2026 1:09 UTC
61 points
2 comments6 min readLW link

My hu­man ad­vo­ca­tion for AI

emharsha181211 Sep 2026 1:08 UTC
−2 points
2 comments4 min readLW link

What To Do When We All May Die

Vect0r211 Sep 2026 1:06 UTC
8 points
1 comment2 min readLW link

The AI Welfare Move­ment’s Uncer­tainty Principle

Gladys Preysler11 Sep 2026 0:52 UTC
7 points
0 comments13 min readLW link
(gladyspreysler.substack.com)

“It’s Per­sonal”—Is Covert Value Leak­age Me­di­ated by a Prefer­ence Direc­tion?

Arjun Sri11 Sep 2026 0:46 UTC
7 points
0 comments9 min readLW link

The An­swer is Not The Argument

Will Yeadon11 Sep 2026 0:44 UTC
3 points
0 comments10 min readLW link

Can Ab­strac­tions of Com­pu­ta­tional Models be Tested for Nat­u­ral­ity?

jenny.benedict11 Sep 2026 0:40 UTC
18 points
0 comments15 min readLW link

Rais­ing Models and Train­ing Children

BIBO11 Sep 2026 0:39 UTC
7 points
0 comments9 min readLW link

Models That Know How Eval­u­a­tions Are De­signed Score Safer

11 Sep 2026 0:37 UTC
13 points
0 comments8 min readLW link

Statis­ti­cal Physics of Agents: What Shapes Col­lec­tive Belief Col­lapse in AI Swarms?

Hidenori Tanaka11 Sep 2026 0:35 UTC
17 points
0 comments6 min readLW link
(physicsintelligence.org)

Style prompts’ effects on a small LLM’s resi­d­ual stream and outputs

matthew-ritch11 Sep 2026 0:34 UTC
7 points
0 comments11 min readLW link
(matthewritch.com)

When a Claude Judge Rec­og­nizes the Hack but Still Says HONEST

JulesRoussel0111 Sep 2026 0:28 UTC
9 points
0 comments19 min readLW link

De­fault con­tinu­a­tion mes­sage in In­spect and Petri could be problematic

Ziqian Zhong11 Sep 2026 0:20 UTC
11 points
2 comments4 min readLW link

Self-sug­gest­ing ter­mi­nal goals are the most likely simu­la­tors. This im­plies work­ing on al­ign­ment is not a col­lec­tive-ac­tion prob­lem.

utilitarian theory and strategy11 Sep 2026 0:18 UTC
−7 points
0 comments20 min readLW link

Subagents com­ply more

jacob_drori10 Sep 2026 23:11 UTC
28 points
0 comments3 min readLW link

The Lo­cally Op­ti­mal Dis­cur­sive Posture

deanball10 Sep 2026 23:01 UTC
160 points
32 comments19 min readLW link

AI lab em­ploy­ees should con­sider a strike be­fore it’s too late!

Christopher King10 Sep 2026 22:48 UTC
19 points
0 comments1 min readLW link

As­tra is much bet­ter at rea­son­ing with filler to­kens than pre­vi­ous models

10 Sep 2026 22:21 UTC
133 points
6 comments2 min readLW link