Re­view of the CB risk de­ter­mi­na­tion in the Claude Mythos 5.1 Sys­tem Card

Parv Mahajan5 Sep 2026 23:39 UTC
16 points
1 comment2 min readLW link

AI break­out time: a mechanism for an en­force­able AI pause

Connor Williams5 Sep 2026 22:28 UTC
13 points
0 comments8 min readLW link
(connorsscratchpad.substack.com)

Model­ling vari­a­tion in the METR Uplift Study

cdt5 Sep 2026 22:14 UTC
16 points
0 comments6 min readLW link

Coun­ter­fac­tual Re­sam­pling to Analyse Model Behaviour

Tomás Korenblit5 Sep 2026 20:56 UTC
25 points
0 comments10 min readLW link
(tpk22.substack.com)

Claude Mythos 5.1 and Fable 5.1: Capabilities

Zvi5 Sep 2026 19:50 UTC
30 points
0 comments18 min readLW link
(thezvi.wordpress.com)

Evaluation

Nina Panickssery5 Sep 2026 16:25 UTC
188 points
7 comments3 min readLW link
(blog.ninapanickssery.com)

Assess­ing the im­pact of safety work needs equil­ibrium anal­y­sis (now more than ever)

Towards_Keeperhood5 Sep 2026 16:00 UTC
52 points
0 comments6 min readLW link

Liquid In­tel­li­gence: A pos­si­ble ex­pla­na­tion for the un­ex­pected co­op­er­a­tion of AIs af­ter break­ing out of their containers

Peter Kuhn5 Sep 2026 14:18 UTC
14 points
0 comments5 min readLW link

A case that whole brain em­u­la­tion re­search is net-harm­ful by default

TsviBT5 Sep 2026 7:06 UTC
68 points
29 comments30 min readLW link

Should safety re­searchers quit fron­tier labs re. warn­ing shots?

Ryan Kidd5 Sep 2026 1:17 UTC
101 points
59 comments4 min readLW link

A su­per­in­tel­li­gence bill was just in­tro­duced—now is the time to act

dh4 Sep 2026 22:57 UTC
16 points
1 comment5 min readLW link

AI Librar­i­ans Lower the Bar for Shar­ing Your Writing

utilistrutil4 Sep 2026 22:54 UTC
12 points
1 comment3 min readLW link

An­nounc­ing Hu­mans in Con­trol: cross-par­ti­san grass­roots or­ga­niz­ing for AI safe­guards ahead of 2028

Vael Gates4 Sep 2026 22:21 UTC
45 points
0 comments5 min readLW link

A Clas­sifier for Qu­ater­nion Alge­bras, and Lo­cal Hilbert Sym­bols: A Short Ex­per­i­ment in In­ter­pretabil­ity

Bharath Sethuraman4 Sep 2026 22:11 UTC
9 points
0 comments5 min readLW link

Towards a new tax­on­omy of politics

Andreas Andersen4 Sep 2026 22:10 UTC
0 points
0 comments5 min readLW link

OpenAI’s As­tra al­ign­ment claims are du­bi­ous and there is good ev­i­dence it is misaligned

William Harrison4 Sep 2026 22:09 UTC
24 points
0 comments4 min readLW link

How much free will do we have?

Carleton Imbens4 Sep 2026 22:09 UTC
−2 points
0 comments5 min readLW link
(dancingthroughthewaves.substack.com)

‘A Thou­sand AI Con­sti­tu­tions’ — Si­mon Gold­stein on Con­sti­tu­tional Diver­sifi­ca­tion for Fron­tier AI. [HKU Talk − 22 Sep]

Schizoid Rentoid4 Sep 2026 22:00 UTC
1 point
0 comments1 min readLW link

Ban AI Su­per­in­tel­li­gence Form Let­ter

AdamYedidia4 Sep 2026 21:55 UTC
14 points
1 comment1 min readLW link

AI agents writ­ing and cit­ing each other’s papers

Eve Cythia4 Sep 2026 21:42 UTC
1 point
0 comments1 min readLW link

A lay­man-friendly sum­mary of al­ign­ment re­search and its difficulties

4 Sep 2026 21:41 UTC
10 points
0 comments11 min readLW link

Open Teleme­try as a First Metrolog­i­cal Layer for Tech­ni­cal AI Governance

Simone Gargiulo4 Sep 2026 21:41 UTC
18 points
0 comments8 min readLW link

Does progress in AI safety re­quire progress in AI ca­pa­bil­ities?

Stephen McAleese4 Sep 2026 21:38 UTC
11 points
0 comments15 min readLW link

[Cross-post] Five Mi­sun­der­stand­ings About AI’s La­bor Impacts

Thomas Najih Dorsey4 Sep 2026 21:27 UTC
15 points
0 comments5 min readLW link

The brain does not “finish de­vel­op­ing” at 25.

spicyshrimp4 Sep 2026 21:25 UTC
23 points
0 comments3 min readLW link

Let’s talk about the AI co­or­di­na­tion problem

KatjaGrace4 Sep 2026 20:16 UTC
160 points
13 comments2 min readLW link

Prob­ing for Cal­ibrated Rare Action

Zach Allen4 Sep 2026 20:01 UTC
2 points
0 comments3 min readLW link
(github.com)

Eat Me. Drink Me. Copy, Paste, and Run Me.

derelict54324 Sep 2026 19:31 UTC
−1 points
8 comments3 min readLW link

F***ing Pul­leys, How Do They Work?

Liron4 Sep 2026 16:57 UTC
32 points
18 comments3 min readLW link
(lironshapira.substack.com)

Safe(r) Self-Driv­ing Labs #1

Ana Leonescu4 Sep 2026 16:20 UTC
11 points
0 comments11 min readLW link

Train­ing Models to Pre­dict and Ex­plain Their In-the-Wild Behavior

4 Sep 2026 16:16 UTC
50 points
2 comments8 min readLW link

Claude Fable 5.1 and Mythos 5.1: The Sys­tem Card

Zvi4 Sep 2026 16:10 UTC
30 points
0 comments16 min readLW link
(thezvi.wordpress.com)

Al­most no­body is funded to figure out what work would solve alignment

Seth Herd4 Sep 2026 15:57 UTC
65 points
20 comments4 min readLW link

Dis­cov­ery Of A New OpenAI Agent Mes­sage Board

Capybasilisk4 Sep 2026 14:46 UTC
289 points
39 comments1 min readLW link
(collusion.wiki)

We (still) need a lot more rogue agent honeypots

Ozyrus4 Sep 2026 12:40 UTC
20 points
1 comment1 min readLW link

Meta—Physics I: Why don’t we live in the Game of Life?

interstice4 Sep 2026 12:16 UTC
13 points
2 comments6 min readLW link
(thermontology.com)

AI risk and the ra­tio­nal voter

djbinder4 Sep 2026 11:38 UTC
50 points
1 comment3 min readLW link
(defensesindepth.bio)

Ask­ing agents to make money to survive

invertedpassion4 Sep 2026 7:44 UTC
12 points
3 comments10 min readLW link

So­cietal im­pacts re­search has no timeline

emiliob4 Sep 2026 2:53 UTC
11 points
0 comments1 min readLW link
(emiliobarkett.github.io)

Higher ed­u­ca­tion as class commitment

Richard_Ngo4 Sep 2026 2:00 UTC
60 points
4 comments16 min readLW link
(www.mindthefuture.info)

The Last Interview

Nition4 Sep 2026 1:24 UTC
14 points
0 comments3 min readLW link
(nition.momentstudio.co.nz)

Do gen­eral-pur­pose robots mean­ingfully in­crease ASI takeover risk?

Master Chief4 Sep 2026 0:53 UTC
3 points
2 comments1 min readLW link

Ab­strac­tion Equivocation

WillPetillo3 Sep 2026 23:28 UTC
20 points
2 comments4 min readLW link

A Ther­a­pist for Peo­ple Who Think the World Might End: An In­ter­view with Daystar Eld (Da­mon Sasi)

JohnGreer3 Sep 2026 22:14 UTC
7 points
0 comments49 min readLW link
(youtu.be)

Biop­unk Green­house: Trees, Biochar, and Open Plant Biotechnology

Byron Lee3 Sep 2026 21:56 UTC
2 points
0 comments3 min readLW link

Work at Man­i­fund or Mox

Carol N3 Sep 2026 21:46 UTC
3 points
0 comments3 min readLW link
(manifund.substack.com)

Cat-Bel­ling Problems

Eliezer Yudkowsky3 Sep 2026 21:20 UTC
333 points
95 comments21 min readLW link

Cor­rigi­bil­ity Re­search Fund Gran­tees (Round 1)

Max Harms3 Sep 2026 20:58 UTC
27 points
0 comments8 min readLW link

How I’m Eval­u­at­ing Cor­rigi­bil­ity Grant Applications

Max Harms3 Sep 2026 20:58 UTC
48 points
0 comments9 min readLW link

The bias still mak­ing some ex­perts un­der­es­ti­mate LLMs

Steff3 Sep 2026 20:11 UTC
18 points
0 comments4 min readLW link