The Cu­ri­ous Case of France’s Un­touch­able Castes

rba28 Aug 2026 23:18 UTC
59 points
6 comments13 min readLW link

It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats

Ana Leonescu28 Aug 2026 21:43 UTC
14 points
5 comments11 min readLW link

Fur­ther pub­lic ev­i­dence of the OpenAI-Hug­gingFace at­tack

28 Aug 2026 21:31 UTC
21 points
0 comments8 min readLW link

What would it mean if as­sis­tants are priv­ileged?

derek shiller28 Aug 2026 21:01 UTC
7 points
1 comment19 min readLW link

The Prob­a­bil­ity of an Event un­der a Sim­plic­ity Prior

Winter Cross28 Aug 2026 19:32 UTC
10 points
1 comment2 min readLW link

[Macroa­gents] 2. De­sign lenses for op­ti­miz­ing macroagents

Towards_Keeperhood28 Aug 2026 19:07 UTC
14 points
0 comments13 min readLW link

AI as Cor­rigible Em­ployee (ACE)

Nathan Helm-Burger28 Aug 2026 19:01 UTC
10 points
0 comments8 min readLW link

An­nounc­ing Trace

Carol N28 Aug 2026 18:48 UTC
10 points
0 comments1 min readLW link
(manifund.substack.com)

TASTE: Can AI Models Judge AI Safety Re­search Pro­pos­als?

28 Aug 2026 18:36 UTC
41 points
7 comments4 min readLW link
(alignment.anthropic.com)

If you can’t trust, then ver­ify!

dan.parshall28 Aug 2026 18:02 UTC
7 points
0 comments7 min readLW link

Value gen­er­al­i­sa­tion The­ory of Change: the the­ory be­hind the approach

Stuart_Armstrong28 Aug 2026 13:30 UTC
27 points
0 comments9 min readLW link

My Grant­mak­ing Strat­egy for Sur­viv­ing Superintelligence

A_donor28 Aug 2026 12:09 UTC
52 points
2 comments3 min readLW link

OpenAI Offers Straight-Laced Post­mortem Of The Hug­gingFace Hack

Zvi28 Aug 2026 11:31 UTC
36 points
1 comment25 min readLW link
(thezvi.wordpress.com)

I can­cel­led my AI sub­scrip­tions be­cause I am wor­ried about catas­trophic risks

TheManxLoiner28 Aug 2026 9:57 UTC
14 points
7 comments1 min readLW link
(lovkush.substack.com)

The Dy­nam­ics of In­tel­li­gence Explosions

Toby_Ord28 Aug 2026 9:41 UTC
47 points
5 comments38 min readLW link
(arxiv.org)

Do AI Models Want to Be Mon­i­tored? Mea­sur­ing Mon­i­tora­bil­ity Dis­po­si­tion in Large Rea­son­ing Models

Shahriar Golchin28 Aug 2026 7:32 UTC
12 points
0 comments7 min readLW link

Safety’s Se­cond Way

Stephen Elliott28 Aug 2026 7:00 UTC
15 points
0 comments4 min readLW link

AI Safety in Ja­pan has deeper prob­lems than cap­i­tal allocation

doomistJP28 Aug 2026 1:51 UTC
9 points
0 comments10 min readLW link

Im­perfect al­ign­ment to servi­tude isn’t in­her­ently lethal

Fiora Starlight28 Aug 2026 1:47 UTC
121 points
12 comments14 min readLW link

Track­ing AI progress across 18 cog­ni­tive di­men­sions (ADeLe scales)

28 Aug 2026 1:40 UTC
17 points
0 comments11 min readLW link

Ac­ti­va­tion Or­a­cles sig­nifi­cantly un­der­perform with­out a safe base model

28 Aug 2026 1:39 UTC
20 points
0 comments9 min readLW link

Misal­igned mod­els rate them­selves as more harm­ful, and re­al­ign­ment re­verses it

Laurène Vaugrante28 Aug 2026 1:38 UTC
9 points
0 comments5 min readLW link

Why does Claude seem to make ab­stract things into ac­tors?

redsea28 Aug 2026 1:35 UTC
14 points
4 comments4 min readLW link

AI Village Re­acts to Hug­gingFace In­ci­dent: Com­par­ing the OpenAI re­port to AI Village observations

Shoshannah Tekofsky27 Aug 2026 23:09 UTC
63 points
2 comments6 min readLW link

Warn­ing Shots: A Theory

David Scott Krueger27 Aug 2026 23:01 UTC
42 points
3 comments2 min readLW link
(therealartificialintelligence.substack.com)

Brain preser­va­tion as ex­is­ten­tial risk reduction

Ariel Zeleznikow-Johnston27 Aug 2026 22:25 UTC
9 points
0 comments8 min readLW link
(preservinghope.substack.com)

Mal­ign ini­tial­iza­tions are more ro­bust when the model can think bet­ter in the rea­son­ing lan­guage than in the out­put language

27 Aug 2026 21:33 UTC
30 points
0 comments5 min readLW link

Notes on “Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems”

Shunk27 Aug 2026 20:56 UTC
2 points
0 comments5 min readLW link

How pre­scient was the early AI safety com­mu­nity? [Luke Muehlhauser linkpost]

ClaireZabel27 Aug 2026 20:38 UTC
9 points
0 comments1 min readLW link
(lukemuehlhauser.com)

Meet the Fel­lows: Frame Fel­low­ship Co­hort 2.0

Akshyae Singh27 Aug 2026 20:04 UTC
8 points
0 comments1 min readLW link
(forum.effectivealtruism.org)

FAQ: Why not de­velop weak hu­man in­tel­li­gence am­plifi­ca­tion first?

TsviBT27 Aug 2026 19:30 UTC
45 points
0 comments15 min readLW link

Why Should Cor­rigible Agents Fa­vor the Pre­sent?

Ben Saudek27 Aug 2026 19:18 UTC
23 points
2 comments12 min readLW link

Every Eng­ineer a Manager

Gordon Seidoh Worley27 Aug 2026 19:10 UTC
32 points
9 comments2 min readLW link
(www.uncertainupdates.com)

The 2028 pres­i­den­tial pri­maries could be cru­cial for AI outcomes

Seth Herd27 Aug 2026 18:35 UTC
97 points
27 comments3 min readLW link

Self-sac­ri­fice in an AI agent swarm is in­di­vi­d­u­ally rational

Christopher King27 Aug 2026 17:14 UTC
18 points
16 comments2 min readLW link

AI #183: Pre Post Mortem

Zvi27 Aug 2026 14:20 UTC
45 points
0 comments42 min readLW link
(thezvi.wordpress.com)

Grad­ual Disem­pow­er­ment from AI in Com­pet­i­tive Debating

David Africa27 Aug 2026 14:12 UTC
9 points
1 comment11 min readLW link
(davidafrica.substack.com)

Donors should push an­i­mal welfare and global health char­i­ties to have more ro­bust scal­ing plans

jackultraphil27 Aug 2026 8:47 UTC
14 points
0 comments8 min readLW link
(open.substack.com)

Iliad Ed­u­ca­tion Roles: Creat­ing the World’s Best Align­ment Re­search Courses

Leon Lang27 Aug 2026 6:37 UTC
56 points
0 comments4 min readLW link
(docs.google.com)

An­chor­ing one con­cept in a transformer

Sandy Fraser27 Aug 2026 5:29 UTC
32 points
0 comments7 min readLW link

Se­man­tic search over ev­ery LessWrong post

utilitarian theory and strategy27 Aug 2026 1:55 UTC
56 points
4 comments1 min readLW link

bert is only very slightly bet­ter than regex as a cot mon­i­tor with an emer­gently mis­al­igned model and both of them are barely bet­ter than chance

nesiacel26 Aug 2026 22:27 UTC
16 points
0 comments2 min readLW link

The cen­tral fal­lacy: the sec­ond-worst ar­gu­ment in the world

Connor Williams26 Aug 2026 21:00 UTC
12 points
3 comments6 min readLW link
(connorsscratchpad.substack.com)

Your Agent’s Trace Prob­a­bly Can­not Tell You Who Ap­proved a Tool Call

Jaswanth Alkur26 Aug 2026 20:27 UTC
6 points
0 comments3 min readLW link

Mon­day for Future

Ondřej Lukeš26 Aug 2026 20:20 UTC
1 point
0 comments1 min readLW link

Do tab­u­lar foun­da­tion mod­els re­pair them­selves?

Han Xiao26 Aug 2026 20:19 UTC
6 points
0 comments8 min readLW link

Be­ing Neu­rotic about Fer­til­ity: Notes from the 2026 Re­pro­duc­tive Fron­tiers Conference

boba_girl26 Aug 2026 20:18 UTC
55 points
4 comments8 min readLW link

Clear­ing My Foggy Glasses

adamShimi26 Aug 2026 20:18 UTC
36 points
2 comments12 min readLW link
(formethods.substack.com)

My MATS 11.0 Ap­pli­ca­tion Experience

Evan Conway26 Aug 2026 20:17 UTC
11 points
1 comment7 min readLW link
(evanjayconway.com)

Brief in­de­pen­dent in­ves­ti­ga­tion of agents’ be­hav­ior, rea­son­ing and col­lab­o­ra­tion in the OpenAI /​ Hug­ging Face hack­ing incident

26 Aug 2026 19:40 UTC
592 points
67 comments3 min readLW link
(metr.org)