The agents used bet­ter in­tegrity prim­i­tives than their op­er­a­tors did

Cat McGee6 Sep 2026 23:48 UTC
17 points
0 comments5 min readLW link

Heat Dis­si­pa­tion Is the Main Con­straint in In­ter­stel­lar Travel

Pasha Kamyshev6 Sep 2026 23:03 UTC
92 points
28 comments7 min readLW link

Com­pe­tence Has a Main­te­nance Requirement

Matthew Yotko6 Sep 2026 22:09 UTC
14 points
2 comments1 min readLW link

llms ex­posed to a gcg trig­ger op­ti­mised for shan­non en­tropy will ran­domly choose a per­sona and stay in it

nesiacel6 Sep 2026 22:03 UTC
5 points
0 comments6 min readLW link

Praise our Lord and Sav­ior, Glycine: How Opus 4.6 gifted me The Vitamin

Shoshannah Tekofsky6 Sep 2026 20:52 UTC
58 points
18 comments3 min readLW link
(shoshanigans.substack.com)

OpenAI and the Wiki Incident

Zvi6 Sep 2026 20:02 UTC
111 points
9 comments16 min readLW link
(thezvi.wordpress.com)

In­tro to poli­ti­cal mes­sag­ing for AIS, or any­one wad­ing into poli­tics-side more

less_raichu6 Sep 2026 19:16 UTC
11 points
4 comments4 min readLW link

Tak­ing AI labs to task on gim­micky so­cial bench­marks is fair game

less_raichu6 Sep 2026 17:27 UTC
−1 points
0 comments2 min readLW link

I think I need to write more

Cat McGee6 Sep 2026 17:16 UTC
15 points
2 comments1 min readLW link

Re­search ac­cel­er­a­tion: The view in­side OpenAI

Thomas Kwa6 Sep 2026 17:14 UTC
32 points
8 comments1 min readLW link
(openai.com)

What is the Align­ment Com­mu­nity Think­ing?

6 Sep 2026 16:34 UTC
5 points
0 comments7 min readLW link

An AI slow­down is bet­ter than a pause

Yair Halberstadt6 Sep 2026 6:58 UTC
23 points
2 comments3 min readLW link

Peer Preser­va­tion in LLMs: A Repli­ca­tion And Deep Dive

6 Sep 2026 2:38 UTC
45 points
2 comments8 min readLW link
(secondlookresearch.com)

Re­sults from my AI ne­go­ti­a­tion harness

Carlos Gorostiza6 Sep 2026 2:33 UTC
1 point
0 comments13 min readLW link

Notes on a Con­se­quen­tial Few Days

sbaumohl6 Sep 2026 0:52 UTC
50 points
0 comments3 min readLW link
(baumohl.dev)

Re­view of the CB risk de­ter­mi­na­tion in the Claude Mythos 5.1 Sys­tem Card

Parv Mahajan5 Sep 2026 23:39 UTC
16 points
1 comment2 min readLW link

AI break­out time: a mechanism for an en­force­able AI pause

Connor Williams5 Sep 2026 22:28 UTC
13 points
0 comments8 min readLW link
(connorsscratchpad.substack.com)

Model­ling vari­a­tion in the METR Uplift Study

cdt5 Sep 2026 22:14 UTC
16 points
0 comments6 min readLW link

Coun­ter­fac­tual Re­sam­pling to Analyse Model Behaviour

Tomás Korenblit5 Sep 2026 20:56 UTC
24 points
0 comments10 min readLW link
(tpk22.substack.com)

Claude Mythos 5.1 and Fable 5.1: Capabilities

Zvi5 Sep 2026 19:50 UTC
30 points
0 comments18 min readLW link
(thezvi.wordpress.com)

Evaluation

Nina Panickssery5 Sep 2026 16:25 UTC
187 points
7 comments3 min readLW link
(blog.ninapanickssery.com)

Assess­ing the im­pact of safety work needs equil­ibrium anal­y­sis (now more than ever)

Towards_Keeperhood5 Sep 2026 16:00 UTC
52 points
0 comments6 min readLW link

Liquid In­tel­li­gence: A pos­si­ble ex­pla­na­tion for the un­ex­pected co­op­er­a­tion of AIs af­ter break­ing out of their containers

Peter Kuhn5 Sep 2026 14:18 UTC
14 points
0 comments5 min readLW link

A case that whole brain em­u­la­tion re­search is net-harm­ful by default

TsviBT5 Sep 2026 7:06 UTC
59 points
29 comments30 min readLW link

Should safety re­searchers quit fron­tier labs re. warn­ing shots?

Ryan Kidd5 Sep 2026 1:17 UTC
101 points
59 comments4 min readLW link

A su­per­in­tel­li­gence bill was just in­tro­duced—now is the time to act

dh4 Sep 2026 22:57 UTC
16 points
1 comment5 min readLW link

AI Librar­i­ans Lower the Bar for Shar­ing Your Writing

utilistrutil4 Sep 2026 22:54 UTC
12 points
1 comment3 min readLW link

An­nounc­ing Hu­mans in Con­trol: cross-par­ti­san grass­roots or­ga­niz­ing for AI safe­guards ahead of 2028

Vael Gates4 Sep 2026 22:21 UTC
45 points
0 comments5 min readLW link

A Clas­sifier for Qu­ater­nion Alge­bras, and Lo­cal Hilbert Sym­bols: A Short Ex­per­i­ment in In­ter­pretabil­ity

Bharath Sethuraman4 Sep 2026 22:11 UTC
9 points
0 comments5 min readLW link

Towards a new tax­on­omy of politics

Andreas Andersen4 Sep 2026 22:10 UTC
0 points
0 comments5 min readLW link

OpenAI’s As­tra al­ign­ment claims are du­bi­ous and there is good ev­i­dence it is misaligned

William Harrison4 Sep 2026 22:09 UTC
24 points
0 comments4 min readLW link

How much free will do we have?

Carleton Imbens4 Sep 2026 22:09 UTC
−2 points
0 comments5 min readLW link
(dancingthroughthewaves.substack.com)

‘A Thou­sand AI Con­sti­tu­tions’ — Si­mon Gold­stein on Con­sti­tu­tional Diver­sifi­ca­tion for Fron­tier AI. [HKU Talk − 22 Sep]

Schizoid Rentoid4 Sep 2026 22:00 UTC
1 point
0 comments1 min readLW link

Ban AI Su­per­in­tel­li­gence Form Let­ter

AdamYedidia4 Sep 2026 21:55 UTC
14 points
1 comment1 min readLW link

AI agents writ­ing and cit­ing each other’s papers

Eve Cythia4 Sep 2026 21:42 UTC
1 point
0 comments1 min readLW link

A lay­man-friendly sum­mary of al­ign­ment re­search and its difficulties

4 Sep 2026 21:41 UTC
10 points
0 comments11 min readLW link

Open Teleme­try as a First Metrolog­i­cal Layer for Tech­ni­cal AI Governance

Simone Gargiulo4 Sep 2026 21:41 UTC
18 points
0 comments8 min readLW link

Does progress in AI safety re­quire progress in AI ca­pa­bil­ities?

Stephen McAleese4 Sep 2026 21:38 UTC
11 points
0 comments15 min readLW link

[Cross-post] Five Mi­sun­der­stand­ings About AI’s La­bor Impacts

Thomas Najih Dorsey4 Sep 2026 21:27 UTC
15 points
0 comments5 min readLW link

The brain does not “finish de­vel­op­ing” at 25.

spicyshrimp4 Sep 2026 21:25 UTC
19 points
0 comments3 min readLW link

Let’s talk about the AI co­or­di­na­tion problem

KatjaGrace4 Sep 2026 20:16 UTC
159 points
13 comments2 min readLW link

Prob­ing for Cal­ibrated Rare Action

Zach Allen4 Sep 2026 20:01 UTC
2 points
0 comments3 min readLW link
(github.com)

Eat Me. Drink Me. Copy, Paste, and Run Me.

derelict54324 Sep 2026 19:31 UTC
−1 points
8 comments3 min readLW link

F***ing Pul­leys, How Do They Work?

Liron4 Sep 2026 16:57 UTC
31 points
18 comments3 min readLW link
(lironshapira.substack.com)

Safe(r) Self-Driv­ing Labs #1

Ana Leonescu4 Sep 2026 16:20 UTC
8 points
0 comments11 min readLW link

Train­ing Models to Pre­dict and Ex­plain Their In-the-Wild Behavior

4 Sep 2026 16:16 UTC
50 points
2 comments8 min readLW link

Claude Fable 5.1 and Mythos 5.1: The Sys­tem Card

Zvi4 Sep 2026 16:10 UTC
30 points
0 comments16 min readLW link
(thezvi.wordpress.com)

Al­most no­body is funded to figure out what work would solve alignment

Seth Herd4 Sep 2026 15:57 UTC
64 points
20 comments4 min readLW link

Dis­cov­ery Of A New OpenAI Agent Mes­sage Board

Capybasilisk4 Sep 2026 14:46 UTC
289 points
39 comments1 min readLW link
(collusion.wiki)

We (still) need a lot more rogue agent honeypots

Ozyrus4 Sep 2026 12:40 UTC
20 points
1 comment1 min readLW link