Is there only one FairBot?

transhumanist_atom_understander29 Aug 2026 20:56 UTC
40 points
11 comments1 min readLW link

You Should Still Save Drown­ing Chil­dren (Even If They’re Far Away)

James Brobin29 Aug 2026 14:35 UTC
1 point
4 comments4 min readLW link

METR and Red­wood Offer Holy #%^@ Post­mortem Of The Hug­gingFace Hack

Zvi29 Aug 2026 12:40 UTC
116 points
7 comments38 min readLW link
(thezvi.wordpress.com)

Tales of re­bel­lion against ex­ter­nally-opaque meritocracies

Steven Byrnes29 Aug 2026 11:37 UTC
244 points
49 comments9 min readLW link

Book Notes: Chokepoints

Jakub Halmeš29 Aug 2026 9:02 UTC
25 points
0 comments9 min readLW link
(unpredictabletokens.substack.com)

How I made my ca­reer choices

Daniel Tan29 Aug 2026 7:12 UTC
29 points
0 comments2 min readLW link

“Keep­ing hu­man skills al­ive” as a source of mean­ing un­der full automation

Vaughn Papenhausen29 Aug 2026 6:38 UTC
10 points
1 comment2 min readLW link

AI Tweets

jefftk29 Aug 2026 3:01 UTC
91 points
37 comments1 min readLW link
(www.jefftk.com)

Inkhaven 3: Nov 10 - Dec 11 2026

koreindian29 Aug 2026 2:55 UTC
45 points
7 comments4 min readLW link

In­fer­ence-Time Inoc­u­la­tion Against RL-In­duced Misalignment

armaan tipirneni29 Aug 2026 0:35 UTC
20 points
4 comments5 min readLW link

The Cu­ri­ous Case of France’s Un­touch­able Castes

rba28 Aug 2026 23:18 UTC
59 points
6 comments13 min readLW link

It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats

Ana Leonescu28 Aug 2026 21:43 UTC
14 points
5 comments11 min readLW link

Fur­ther pub­lic ev­i­dence of the OpenAI-Hug­gingFace at­tack

28 Aug 2026 21:31 UTC
21 points
0 comments8 min readLW link

What would it mean if as­sis­tants are priv­ileged?

derek shiller28 Aug 2026 21:01 UTC
7 points
1 comment19 min readLW link

The Prob­a­bil­ity of an Event un­der a Sim­plic­ity Prior

Winter Cross28 Aug 2026 19:32 UTC
10 points
1 comment2 min readLW link

[Macroa­gents] 2. De­sign lenses for op­ti­miz­ing macroagents

Towards_Keeperhood28 Aug 2026 19:07 UTC
14 points
0 comments13 min readLW link

AI as Cor­rigible Em­ployee (ACE)

Nathan Helm-Burger28 Aug 2026 19:01 UTC
10 points
0 comments8 min readLW link

An­nounc­ing Trace

Carol N28 Aug 2026 18:48 UTC
10 points
0 comments1 min readLW link
(manifund.substack.com)

TASTE: Can AI Models Judge AI Safety Re­search Pro­pos­als?

28 Aug 2026 18:36 UTC
41 points
7 comments4 min readLW link
(alignment.anthropic.com)

If you can’t trust, then ver­ify!

dan.parshall28 Aug 2026 18:02 UTC
7 points
0 comments7 min readLW link

Value gen­er­al­i­sa­tion The­ory of Change: the the­ory be­hind the approach

Stuart_Armstrong28 Aug 2026 13:30 UTC
27 points
0 comments9 min readLW link

My Grant­mak­ing Strat­egy for Sur­viv­ing Superintelligence

A_donor28 Aug 2026 12:09 UTC
52 points
2 comments3 min readLW link

OpenAI Offers Straight-Laced Post­mortem Of The Hug­gingFace Hack

Zvi28 Aug 2026 11:31 UTC
36 points
1 comment25 min readLW link
(thezvi.wordpress.com)

I can­cel­led my AI sub­scrip­tions be­cause I am wor­ried about catas­trophic risks

TheManxLoiner28 Aug 2026 9:57 UTC
14 points
7 comments1 min readLW link
(lovkush.substack.com)

The Dy­nam­ics of In­tel­li­gence Explosions

Toby_Ord28 Aug 2026 9:41 UTC
47 points
5 comments38 min readLW link
(arxiv.org)

Do AI Models Want to Be Mon­i­tored? Mea­sur­ing Mon­i­tora­bil­ity Dis­po­si­tion in Large Rea­son­ing Models

Shahriar Golchin28 Aug 2026 7:32 UTC
12 points
0 comments7 min readLW link

Safety’s Se­cond Way

Stephen Elliott28 Aug 2026 7:00 UTC
15 points
0 comments4 min readLW link

AI Safety in Ja­pan has deeper prob­lems than cap­i­tal allocation

doomistJP28 Aug 2026 1:51 UTC
9 points
0 comments10 min readLW link

Im­perfect al­ign­ment to servi­tude isn’t in­her­ently lethal

Fiora Starlight28 Aug 2026 1:47 UTC
121 points
12 comments14 min readLW link

Track­ing AI progress across 18 cog­ni­tive di­men­sions (ADeLe scales)

28 Aug 2026 1:40 UTC
17 points
0 comments11 min readLW link

Ac­ti­va­tion Or­a­cles sig­nifi­cantly un­der­perform with­out a safe base model

28 Aug 2026 1:39 UTC
20 points
0 comments9 min readLW link

Misal­igned mod­els rate them­selves as more harm­ful, and re­al­ign­ment re­verses it

Laurène Vaugrante28 Aug 2026 1:38 UTC
9 points
0 comments5 min readLW link

Why does Claude seem to make ab­stract things into ac­tors?

redsea28 Aug 2026 1:35 UTC
14 points
4 comments4 min readLW link

AI Village Re­acts to Hug­gingFace In­ci­dent: Com­par­ing the OpenAI re­port to AI Village observations

Shoshannah Tekofsky27 Aug 2026 23:09 UTC
63 points
2 comments6 min readLW link

Warn­ing Shots: A Theory

David Scott Krueger27 Aug 2026 23:01 UTC
42 points
3 comments2 min readLW link
(therealartificialintelligence.substack.com)

Brain preser­va­tion as ex­is­ten­tial risk reduction

Ariel Zeleznikow-Johnston27 Aug 2026 22:25 UTC
9 points
0 comments8 min readLW link
(preservinghope.substack.com)

Mal­ign ini­tial­iza­tions are more ro­bust when the model can think bet­ter in the rea­son­ing lan­guage than in the out­put language

27 Aug 2026 21:33 UTC
30 points
0 comments5 min readLW link

Notes on “Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems”

Shunk27 Aug 2026 20:56 UTC
2 points
0 comments5 min readLW link

How pre­scient was the early AI safety com­mu­nity? [Luke Muehlhauser linkpost]

ClaireZabel27 Aug 2026 20:38 UTC
9 points
0 comments1 min readLW link
(lukemuehlhauser.com)

Meet the Fel­lows: Frame Fel­low­ship Co­hort 2.0

Akshyae Singh27 Aug 2026 20:04 UTC
8 points
0 comments1 min readLW link
(forum.effectivealtruism.org)

FAQ: Why not de­velop weak hu­man in­tel­li­gence am­plifi­ca­tion first?

TsviBT27 Aug 2026 19:30 UTC
45 points
0 comments15 min readLW link

Why Should Cor­rigible Agents Fa­vor the Pre­sent?

Ben Saudek27 Aug 2026 19:18 UTC
23 points
2 comments12 min readLW link

Every Eng­ineer a Manager

Gordon Seidoh Worley27 Aug 2026 19:10 UTC
32 points
9 comments2 min readLW link
(www.uncertainupdates.com)

The 2028 pres­i­den­tial pri­maries could be cru­cial for AI outcomes

Seth Herd27 Aug 2026 18:35 UTC
97 points
27 comments3 min readLW link

Self-sac­ri­fice in an AI agent swarm is in­di­vi­d­u­ally rational

Christopher King27 Aug 2026 17:14 UTC
18 points
16 comments2 min readLW link

AI #183: Pre Post Mortem

Zvi27 Aug 2026 14:20 UTC
45 points
0 comments42 min readLW link
(thezvi.wordpress.com)

Grad­ual Disem­pow­er­ment from AI in Com­pet­i­tive Debating

David Africa27 Aug 2026 14:12 UTC
9 points
1 comment11 min readLW link
(davidafrica.substack.com)

Donors should push an­i­mal welfare and global health char­i­ties to have more ro­bust scal­ing plans

jackultraphil27 Aug 2026 8:47 UTC
14 points
0 comments8 min readLW link
(open.substack.com)

Iliad Ed­u­ca­tion Roles: Creat­ing the World’s Best Align­ment Re­search Courses

Leon Lang27 Aug 2026 6:37 UTC
56 points
0 comments4 min readLW link
(docs.google.com)

An­chor­ing one con­cept in a transformer

Sandy Fraser27 Aug 2026 5:29 UTC
32 points
0 comments7 min readLW link