On­line Se­quences Book Club: Begin­ners Wel­come!

gluesniffer198410 Aug 2026 22:10 UTC
1 point
0 comments1 min readLW link

The Pac­ing of the Frontier

Zvi10 Aug 2026 21:50 UTC
28 points
0 comments19 min readLW link
(thezvi.wordpress.com)

Q: Is dual-use an al­ign­ment-com­plete prob­lem?

kapedalex10 Aug 2026 21:40 UTC
10 points
1 comment1 min readLW link

Does post-train­ing quan­ti­za­tion change welfare-rele­vant in­di­ca­tors in open-weight lan­guage mod­els?

ashesfall10 Aug 2026 21:09 UTC
9 points
0 comments8 min readLW link

Claude sum­ma­rizes be­hav­ior as sig­nifi­cantly less mis­al­igned when the ac­tor is Claude vs an­other model

Ezra Newman10 Aug 2026 17:20 UTC
58 points
8 comments1 min readLW link

Off-policy hon­esty train­ing gen­er­al­izes bet­ter than on-policy hon­esty training

10 Aug 2026 16:31 UTC
20 points
0 comments9 min readLW link

Four LLM loss func­tions → four fla­vors of LLM misalignment

Steven Byrnes10 Aug 2026 16:16 UTC
383 points
31 comments6 min readLW link

Co­er­cion and De­cep­tion in AI-to-AI Management

10 Aug 2026 16:13 UTC
12 points
0 comments8 min readLW link
(compassionalignedml.substack.com)

You’re Ab­solutely Right

Linch10 Aug 2026 16:04 UTC
183 points
5 comments8 min readLW link
(linch.substack.com)

Book Re­view: The In­finity Machine

Mikewins10 Aug 2026 16:01 UTC
10 points
0 comments1 min readLW link

On De­moc­ra­tiz­ing ASI to Pre­serve Civil Liberties

MichaelDickens10 Aug 2026 12:08 UTC
43 points
5 comments3 min readLW link

Is Eval Gam­ing Down­stream of Ver­bal­ized Eval Aware­ness? Not when it’s re­flex­ive.

Kieron Kretschmar10 Aug 2026 8:24 UTC
14 points
2 comments11 min readLW link

Is it eth­i­cal to work on gen­eral-pur­pose robots given the risk of to­tal­i­tar­i­anism?

Master Chief10 Aug 2026 6:40 UTC
22 points
5 comments1 min readLW link

How to be an AI safety re­search engineer

Ruben Castaing10 Aug 2026 5:32 UTC
12 points
1 comment6 min readLW link

Hiring Vibe-wran­gler Match­mak­ing Thread

Raemon10 Aug 2026 3:53 UTC
26 points
1 comment3 min readLW link

How to get an­swers to ques­tions that con­fuse you (maybe)

Elijah10 Aug 2026 1:35 UTC
4 points
0 comments5 min readLW link

The Agen­tic Clusterfuck

Chapin Lenthall-Cleary10 Aug 2026 1:27 UTC
51 points
15 comments1 min readLW link

Over­think­ing: Am­plify­ing rea­son­ing weights makes mod­els re­veal their secrets

9 Aug 2026 22:06 UTC
22 points
0 comments7 min readLW link
(arxiv.org)

AI-am­plified demo­cratic back­slid­ing: an exploration

9 Aug 2026 22:05 UTC
17 points
3 comments6 min readLW link

The Apoca­lyp­tic Ar­rival of Truth

Caleb Biddulph9 Aug 2026 21:32 UTC
43 points
2 comments1 min readLW link

“Com­mu­nity Notes” re­s­olu­tion for vague predictions

Raemon9 Aug 2026 19:51 UTC
43 points
22 comments1 min readLW link

Ten Thou­sand Cy­ber Labs for Train­ing & Eval

TheVinci9 Aug 2026 17:57 UTC
10 points
3 comments3 min readLW link

A challenge: Can you make an LLM fol­low these in­struc­tions?

Steff9 Aug 2026 16:50 UTC
18 points
7 comments6 min readLW link

What just hap­pened? A ret­ro­spec­tive of AI alignment

Richard_Ngo9 Aug 2026 15:58 UTC
640 points
152 comments16 min readLW link

Who does the con­fess­ing, and will they con­fess to anything

Abhishek Mishra9 Aug 2026 13:34 UTC
7 points
0 comments8 min readLW link

The world will be full of “sci-fi” things, and ev­ery­one will be unim­pressed and disappointed

Expertium9 Aug 2026 13:10 UTC
88 points
21 comments3 min readLW link

A Spillway for Agent Coordination

Kaustubh Kislay9 Aug 2026 6:51 UTC
30 points
0 comments6 min readLW link

Canobie Lake Visit

jefftk9 Aug 2026 2:30 UTC
12 points
0 comments4 min readLW link
(www.jefftk.com)

Dutch-book re­sis­tant prob­a­bil­ity over cen­tered worlds

jessicata8 Aug 2026 23:05 UTC
31 points
7 comments7 min readLW link
(unstableontology.com)

Why Low Fer­til­ity Rates Are a Pos­i­tive Feed­back Loop

LoopGameScrollMonkey8 Aug 2026 20:17 UTC
46 points
1 comment25 min readLW link

Glimpses of superintelligence

PratyushRT8 Aug 2026 20:15 UTC
7 points
0 comments3 min readLW link

‘AI Es­caped Its Sand­box’ — What Does That Ac­tu­ally Mean?

Jakub Halmeš8 Aug 2026 18:55 UTC
13 points
2 comments6 min readLW link
(unpredictabletokens.substack.com)

What Hap­pened: OpenAI and HuggingFace

Zvi8 Aug 2026 16:10 UTC
111 points
7 comments11 min readLW link
(thezvi.wordpress.com)

In­duc­ing self-other over­lap with SFT re­duces de­cep­tion at scale, but gen­er­al­iza­tion re­mains uneven

Marc Carauleanu8 Aug 2026 15:21 UTC
24 points
1 comment7 min readLW link

FAQ: Isn’t AGI com­ing too soon for re­pro­ge­net­ics to help?

TsviBT8 Aug 2026 8:54 UTC
150 points
69 comments14 min readLW link

Rea­son­ing was not made for Deduction

epicurus8 Aug 2026 8:42 UTC
15 points
1 comment11 min readLW link

Sim­plify­ing the an­thropic im­pos­si­bil­ity result

Stuart_Armstrong8 Aug 2026 8:25 UTC
24 points
13 comments3 min readLW link

CLT Fea­tures Sharpen the Cycli­cal Day-of-Week Man­i­fold in Gemma-2-2b

Anna Marbut8 Aug 2026 2:30 UTC
7 points
0 comments5 min readLW link

AI Reg­u­la­tion Map: a view of AI gov­er­nance in 196 countries

Ria Deane8 Aug 2026 1:23 UTC
7 points
0 comments3 min readLW link

Don’t Build Mindreading

KellerScholl8 Aug 2026 0:30 UTC
298 points
69 comments4 min readLW link
(keller.substack.com)

Self-mon­i­tor­ing doesn’t scale (with­out these 3 coun­ter­mea­sures)

Morgan S7 Aug 2026 21:39 UTC
23 points
0 comments5 min readLW link

Public ev­i­dence of the OpenAI-Hug­gingFace AI attack

beyarkay (Boyd Kane)7 Aug 2026 21:21 UTC
94 points
6 comments7 min readLW link

Don’t Inoc­u­late Every­thing: Strat­ified Inoc­u­la­tion Prompt­ing Nar­rows Back­doors and Pre­serves De­sired Traits

7 Aug 2026 21:16 UTC
45 points
3 comments13 min readLW link

Job-Less Utopia: Macroe­co­nomics in the Age of AGI

Marcus Hutter7 Aug 2026 20:15 UTC
35 points
3 comments1 min readLW link

AI Safety at the Fron­tier: Paper High­lights of July 2026

gasteigerjo7 Aug 2026 19:50 UTC
5 points
0 comments1 min readLW link

Con­tra MacAskill on sav­ing money for the in­tel­li­gence explosion

Carol N7 Aug 2026 17:30 UTC
17 points
0 comments2 min readLW link
(manifund.substack.com)

How to pace the US frontier

7 Aug 2026 17:19 UTC
98 points
1 comment21 min readLW link
(blog.aifutures.org)

OpenAI Trained Its Models For Months While Those Models Were Co­or­di­nat­ing Ex­ploits Via Mes­sage Boards

Zvi7 Aug 2026 17:01 UTC
164 points
15 comments41 min readLW link
(thezvi.wordpress.com)

Item Re­sponse The­ory for AI Safety

7 Aug 2026 16:59 UTC
22 points
2 comments7 min readLW link

Some wild meta­physics that seem­ingly ev­ery­one needs to ac­cept?

Elias Schmied7 Aug 2026 16:13 UTC
16 points
5 comments6 min readLW link