There Should Be More AI Safety Hubs

Seth Lifland8 Jul 2026 23:09 UTC
13 points
1 comment3 min readLW link

Models are blind out­side the J-space. NLAs aren’t.

Pranav Viswanath8 Jul 2026 23:09 UTC
28 points
4 comments9 min readLW link

Solv­ing the BlueDot Puz­zle TAIS: The Ve­loc­ity Ring

Karine Levonyan8 Jul 2026 23:03 UTC
8 points
0 comments4 min readLW link
(karinelevonyan.github.io)

Ac­tion as Choice Ex­pressed Through Move­ment Toward a Goal: a Frame­work for Over­com­ing In­ac­tion

techandsundry8 Jul 2026 23:02 UTC
7 points
0 comments4 min readLW link

Op­ti­mum num­ber of items to in­spect be­fore buy­ing one

bilibili8 Jul 2026 21:23 UTC
18 points
0 comments4 min readLW link

Free will as a model parameter

darshanav8 Jul 2026 21:11 UTC
10 points
1 comment6 min readLW link

Find fund­ing, fast

Austin Chen8 Jul 2026 21:10 UTC
43 points
0 comments3 min readLW link
(manifund.substack.com)

Child­hood and Ed­u­ca­tion #20: Phones and Screens

Zvi8 Jul 2026 19:41 UTC
37 points
1 comment17 min readLW link
(thezvi.wordpress.com)

AI Safety Hong Kong Read­ing Group: Ma­chine Learn­ing and Hu­man Values

Schizoid Rentoid8 Jul 2026 19:19 UTC
1 point
0 comments1 min readLW link

Can the U.S. and China Deny AI?

8 Jul 2026 18:45 UTC
20 points
0 comments14 min readLW link
(secondstrike.substack.com)

A Real-Life Ex­am­ple of an Aligned Sys­tem Killing Hun­dreds of People

London L.8 Jul 2026 18:43 UTC
3 points
8 comments4 min readLW link

Notes on tech­ni­cal al­ign­ment via hu­man-like so­cial drives

Steven Byrnes8 Jul 2026 18:30 UTC
74 points
17 comments42 min readLW link

Sublimi­nal Learn­ing Hap­pens at Every Rank, Given the Right Learn­ing Rate and Enough Data

Lawrence Feng8 Jul 2026 17:32 UTC
60 points
5 comments7 min readLW link

AI Safety at the Fron­tier: Paper High­lights of May & June 2026

gasteigerjo8 Jul 2026 17:19 UTC
16 points
0 comments10 min readLW link

Why study proto-train­ing gam­ing as an ad­ver­sar­ial al­ign­ment failure mode?

8 Jul 2026 17:07 UTC
55 points
0 comments7 min readLW link

Why study al­ign­ment in­ter­ven­tions on pre-RL check­points?

8 Jul 2026 17:07 UTC
66 points
2 comments6 min readLW link

Ar­tifi­cial Am­bas­sadors: The Three Cases for Send­ing LLMs to Deep Space

Ekaterina Matson8 Jul 2026 16:08 UTC
7 points
0 comments12 min readLW link
(medium.com)

Refram­ing LessWrong-style de­ci­sion the­ory as “com­mit­ment the­ory”

Elias Schmied8 Jul 2026 15:25 UTC
61 points
33 comments11 min readLW link

Why I’m a moral anti-re­al­ist but may be un­able to con­vince you

Nina Panickssery8 Jul 2026 15:23 UTC
23 points
12 comments7 min readLW link
(blog.ninapanickssery.com)

AI Fore­cast­ing in 2026: What 11 Analy­ses Say

Ben Wilson8 Jul 2026 14:39 UTC
25 points
0 comments18 min readLW link
(www.metaculus.com)

Why mod­els game evals might mat­ter as much as whether they do it

Igor Ivanov8 Jul 2026 14:30 UTC
30 points
8 comments4 min readLW link

The Em­pa­thetic Poor

a unemployed pastor- de S Brito Gabriel8 Jul 2026 14:21 UTC
5 points
2 comments16 min readLW link

How Valuable are Tra­di­tions?

martinkunev8 Jul 2026 13:44 UTC
10 points
5 comments5 min readLW link

How much slower does take­off go with 10× less com­pute?

brendanhalstead8 Jul 2026 10:04 UTC
46 points
4 comments4 min readLW link

Hu­man Em­pow­er­ment in an AI Society

Alvin Ånestrand8 Jul 2026 9:01 UTC
24 points
0 comments36 min readLW link

The mosquito bucket of doom works

dominicq8 Jul 2026 8:18 UTC
135 points
10 comments5 min readLW link
(blog.d11r.eu)

[Linkpost]Six Story Prompts I Want to Read

Linch8 Jul 2026 6:12 UTC
15 points
0 comments11 min readLW link
(linch.substack.com)

the al­ls­ton lab rat ex­is­ten­tial framing

somayya8 Jul 2026 2:45 UTC
−3 points
0 comments2 min readLW link

the poly­se­man­tic­ity of poly­se­man­tic­ity in lan­guage models

Ayesha Imran8 Jul 2026 2:24 UTC
8 points
0 comments4 min readLW link

Balanc­ing Ri­gor and Utility: A Re­view of “A Prag­matic Vi­sion for In­ter­pretabil­ity”

Sohib Ibrahim8 Jul 2026 2:23 UTC
10 points
0 comments6 min readLW link
(www.lesswrong.com)

How did we get to democ­racy?

8 Jul 2026 2:21 UTC
11 points
0 comments12 min readLW link

[Question] Did OpenClaw cause an up­date to your pri­ors? Have there been other such mo­ments since?

ihatenumbersinusernames78 Jul 2026 0:38 UTC
12 points
3 comments1 min readLW link

No Space Like J-Space

Zvi7 Jul 2026 21:50 UTC
51 points
2 comments18 min readLW link
(thezvi.wordpress.com)

Open-source LLMs ad­minister max­i­mum elec­tric shocks in a Mil­gram-like obe­di­ence experiment

7 Jul 2026 20:05 UTC
10 points
0 comments25 min readLW link
(arxiv.org)

“Fun­da­men­tal Uncer­tainty” Au­dio­book Available

Gordon Seidoh Worley7 Jul 2026 20:00 UTC
15 points
0 comments1 min readLW link
(www.uncertainupdates.com)

Per­sonascope: Mea­sur­ing how deeply LLMs adopt personas

7 Jul 2026 18:38 UTC
42 points
7 comments19 min readLW link

Su­per­hu­man Ar­tic­u­lacy as an LLM Safety Target

Dylan Bowman7 Jul 2026 18:29 UTC
53 points
9 comments5 min readLW link

Cal­ibrat­ing al­ign­ment evals

darshanav7 Jul 2026 18:24 UTC
9 points
0 comments6 min readLW link

June 2026 Links

nomagicpill7 Jul 2026 18:23 UTC
10 points
0 comments5 min readLW link
(nomagicpill.substack.com)

Prob­ing is not enough; a val­idity au­dit for any probe

Ratnaditya J7 Jul 2026 18:10 UTC
7 points
0 comments10 min readLW link

Try­ing to grok An­thropic’s Global Workspace pa­per and the J-space

TheManxLoiner7 Jul 2026 14:52 UTC
14 points
1 comment3 min readLW link

En­tan­gle­ment Between an AI and Its Environment

7 Jul 2026 13:28 UTC
27 points
2 comments11 min readLW link
(limits-of-evaluation.org)

A con­cep­tor by any other name

Keenan Pepper7 Jul 2026 10:09 UTC
60 points
1 comment9 min readLW link

Another Look at ‘Slack’

AdamPiovarchy7 Jul 2026 4:55 UTC
14 points
0 comments5 min readLW link

Join Our AI Safety × Philos­o­phy Read­ing Group

Ahmed7 Jul 2026 4:55 UTC
3 points
0 comments1 min readLW link

The Geom­e­try of Yes: Map­ping Sy­co­phancy In­side an LLM’s Emo­tion Space

Pushpita Das7 Jul 2026 4:54 UTC
11 points
0 comments10 min readLW link

AI Safety Can’t Afford a Se­cond Cause

atlasaligned7 Jul 2026 4:48 UTC
37 points
9 comments3 min readLW link

Ranges of Prob­a­bil­ities: What Are They For?

SaltAndMetal7 Jul 2026 4:47 UTC
9 points
3 comments7 min readLW link

Ar­chi­tec­ture mat­ters for multi-agent security

7 Jul 2026 4:46 UTC
21 points
0 comments10 min readLW link

Data fil­ter­ing works a lot worse than you would ex­pect

7 Jul 2026 4:41 UTC
59 points
9 comments3 min readLW link