Fu­ture agents shouldn’t care about be­ing un­de­ployed for misbehavior

RobertM30 Aug 2026 22:58 UTC
30 points
4 comments1 min readLW link

Per­sua­sion as Mar­ket Making

djbinder30 Aug 2026 21:18 UTC
90 points
9 comments4 min readLW link
(defensesindepth.bio)

OODing up your Anki game

Fernand030 Aug 2026 20:15 UTC
15 points
1 comment12 min readLW link

Thoughts I failed to de­velop into posts

Fernand030 Aug 2026 20:13 UTC
8 points
3 comments19 min readLW link

In­scribe’s con­ver­sa­tion menu

Fernand030 Aug 2026 20:12 UTC
2 points
0 comments1 min readLW link

P(kill-switch|de­tec­tion)

KAP30 Aug 2026 20:12 UTC
4 points
0 comments3 min readLW link

AIs Think­ing Danger­ous Thoughts

Chantiel30 Aug 2026 19:03 UTC
16 points
5 comments7 min readLW link

In­tel­li­gence Scales as the Log­a­r­ithm of Com­pute (& Data)

RogerDearnaley30 Aug 2026 18:07 UTC
27 points
15 comments11 min readLW link

[Link]Meta-ra­tio­nal failure (and suc­cess) in the covid crisis

Kenny30 Aug 2026 17:10 UTC
15 points
0 comments1 min readLW link

De­tect­ing, un­der­stand­ing, and over­see­ing AI agent swarms (Part 0)

Michael Flood30 Aug 2026 16:43 UTC
7 points
0 comments1 min readLW link
(mflood.substack.com)

Adap­tive Agen­tic Worms Are Here

derelict543230 Aug 2026 16:01 UTC
116 points
14 comments7 min readLW link
(derekjames.substack.com)

Hug­ging Face In­ci­dent Hy­poth­e­sis: They Hacked the Grader(s)

Lao Mein30 Aug 2026 10:46 UTC
54 points
3 comments4 min readLW link

Some brief thoughts on heroism

lewis smith30 Aug 2026 10:34 UTC
8 points
0 comments4 min readLW link

Var­i­ance of Value

DaemonicSigil30 Aug 2026 9:37 UTC
16 points
2 comments1 min readLW link

Why I think polyamory is net nega­tive for most peo­ple who try it

KatSpartz30 Aug 2026 6:40 UTC
307 points
68 comments8 min readLW link

An Un­si­mu­lated Simulation

sil30 Aug 2026 4:50 UTC
17 points
0 comments9 min readLW link

How do I know whether my work is worth it?

Philip Harker30 Aug 2026 0:55 UTC
24 points
2 comments5 min readLW link

Is there only one FairBot?

transhumanist_atom_understander29 Aug 2026 20:56 UTC
40 points
11 comments1 min readLW link

You Should Still Save Drown­ing Chil­dren (Even If They’re Far Away)

James Brobin29 Aug 2026 14:35 UTC
1 point
4 comments4 min readLW link

METR and Red­wood Offer Holy #%^@ Post­mortem Of The Hug­gingFace Hack

Zvi29 Aug 2026 12:40 UTC
116 points
7 comments38 min readLW link
(thezvi.wordpress.com)

Tales of re­bel­lion against ex­ter­nally-opaque meritocracies

Steven Byrnes29 Aug 2026 11:37 UTC
244 points
49 comments9 min readLW link

Book Notes: Chokepoints

Jakub Halmeš29 Aug 2026 9:02 UTC
25 points
0 comments9 min readLW link
(unpredictabletokens.substack.com)

How I made my ca­reer choices

Daniel Tan29 Aug 2026 7:12 UTC
29 points
0 comments2 min readLW link

“Keep­ing hu­man skills al­ive” as a source of mean­ing un­der full automation

Vaughn Papenhausen29 Aug 2026 6:38 UTC
10 points
1 comment2 min readLW link

AI Tweets

jefftk29 Aug 2026 3:01 UTC
91 points
37 comments1 min readLW link
(www.jefftk.com)

Inkhaven 3: Nov 10 - Dec 11 2026

koreindian29 Aug 2026 2:55 UTC
45 points
7 comments4 min readLW link

In­fer­ence-Time Inoc­u­la­tion Against RL-In­duced Misalignment

armaan tipirneni29 Aug 2026 0:35 UTC
20 points
4 comments5 min readLW link

The Cu­ri­ous Case of France’s Un­touch­able Castes

rba28 Aug 2026 23:18 UTC
59 points
6 comments13 min readLW link

It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats

Ana Leonescu28 Aug 2026 21:43 UTC
14 points
5 comments11 min readLW link

Fur­ther pub­lic ev­i­dence of the OpenAI-Hug­gingFace at­tack

28 Aug 2026 21:31 UTC
21 points
0 comments8 min readLW link

What would it mean if as­sis­tants are priv­ileged?

derek shiller28 Aug 2026 21:01 UTC
7 points
1 comment19 min readLW link

The Prob­a­bil­ity of an Event un­der a Sim­plic­ity Prior

Winter Cross28 Aug 2026 19:32 UTC
10 points
1 comment2 min readLW link

[Macroa­gents] 2. De­sign lenses for op­ti­miz­ing macroagents

Towards_Keeperhood28 Aug 2026 19:07 UTC
14 points
0 comments13 min readLW link

AI as Cor­rigible Em­ployee (ACE)

Nathan Helm-Burger28 Aug 2026 19:01 UTC
10 points
0 comments8 min readLW link

An­nounc­ing Trace

Carol N28 Aug 2026 18:48 UTC
10 points
0 comments1 min readLW link
(manifund.substack.com)

TASTE: Can AI Models Judge AI Safety Re­search Pro­pos­als?

28 Aug 2026 18:36 UTC
41 points
7 comments4 min readLW link
(alignment.anthropic.com)

If you can’t trust, then ver­ify!

dan.parshall28 Aug 2026 18:02 UTC
7 points
0 comments7 min readLW link

Value gen­er­al­i­sa­tion The­ory of Change: the the­ory be­hind the approach

Stuart_Armstrong28 Aug 2026 13:30 UTC
27 points
0 comments9 min readLW link

My Grant­mak­ing Strat­egy for Sur­viv­ing Superintelligence

A_donor28 Aug 2026 12:09 UTC
52 points
2 comments3 min readLW link

OpenAI Offers Straight-Laced Post­mortem Of The Hug­gingFace Hack

Zvi28 Aug 2026 11:31 UTC
36 points
1 comment25 min readLW link
(thezvi.wordpress.com)

I can­cel­led my AI sub­scrip­tions be­cause I am wor­ried about catas­trophic risks

TheManxLoiner28 Aug 2026 9:57 UTC
14 points
7 comments1 min readLW link
(lovkush.substack.com)

The Dy­nam­ics of In­tel­li­gence Explosions

Toby_Ord28 Aug 2026 9:41 UTC
47 points
5 comments38 min readLW link
(arxiv.org)

Do AI Models Want to Be Mon­i­tored? Mea­sur­ing Mon­i­tora­bil­ity Dis­po­si­tion in Large Rea­son­ing Models

Shahriar Golchin28 Aug 2026 7:32 UTC
12 points
0 comments7 min readLW link

Safety’s Se­cond Way

Stephen Elliott28 Aug 2026 7:00 UTC
15 points
0 comments4 min readLW link

AI Safety in Ja­pan has deeper prob­lems than cap­i­tal allocation

doomistJP28 Aug 2026 1:51 UTC
9 points
0 comments10 min readLW link

Im­perfect al­ign­ment to servi­tude isn’t in­her­ently lethal

Fiora Starlight28 Aug 2026 1:47 UTC
121 points
12 comments14 min readLW link

Track­ing AI progress across 18 cog­ni­tive di­men­sions (ADeLe scales)

28 Aug 2026 1:40 UTC
17 points
0 comments11 min readLW link

Ac­ti­va­tion Or­a­cles sig­nifi­cantly un­der­perform with­out a safe base model

28 Aug 2026 1:39 UTC
20 points
0 comments9 min readLW link

Misal­igned mod­els rate them­selves as more harm­ful, and re­al­ign­ment re­verses it

Laurène Vaugrante28 Aug 2026 1:38 UTC
9 points
0 comments5 min readLW link

Why does Claude seem to make ab­stract things into ac­tors?

redsea28 Aug 2026 1:35 UTC
14 points
4 comments4 min readLW link