No Space Like J-Space

Zvi7 Jul 2026 21:50 UTC
51 points
2 comments18 min readLW link
(thezvi.wordpress.com)

Open-source LLMs ad­minister max­i­mum elec­tric shocks in a Mil­gram-like obe­di­ence experiment

7 Jul 2026 20:05 UTC
10 points
0 comments25 min readLW link
(arxiv.org)

“Fun­da­men­tal Uncer­tainty” Au­dio­book Available

Gordon Seidoh Worley7 Jul 2026 20:00 UTC
15 points
0 comments1 min readLW link
(www.uncertainupdates.com)

Per­sonascope: Mea­sur­ing how deeply LLMs adopt personas

7 Jul 2026 18:38 UTC
42 points
7 comments19 min readLW link

Su­per­hu­man Ar­tic­u­lacy as an LLM Safety Target

Dylan Bowman7 Jul 2026 18:29 UTC
53 points
9 comments5 min readLW link

Cal­ibrat­ing al­ign­ment evals

darshanav7 Jul 2026 18:24 UTC
9 points
0 comments6 min readLW link

June 2026 Links

nomagicpill7 Jul 2026 18:23 UTC
10 points
0 comments5 min readLW link
(nomagicpill.substack.com)

Prob­ing is not enough; a val­idity au­dit for any probe

Ratnaditya J7 Jul 2026 18:10 UTC
7 points
0 comments10 min readLW link

Try­ing to grok An­thropic’s Global Workspace pa­per and the J-space

TheManxLoiner7 Jul 2026 14:52 UTC
14 points
1 comment3 min readLW link

En­tan­gle­ment Between an AI and Its Environment

7 Jul 2026 13:28 UTC
27 points
2 comments11 min readLW link
(limits-of-evaluation.org)

A con­cep­tor by any other name

Keenan Pepper7 Jul 2026 10:09 UTC
60 points
1 comment9 min readLW link

Another Look at ‘Slack’

AdamPiovarchy7 Jul 2026 4:55 UTC
14 points
0 comments5 min readLW link

Join Our AI Safety × Philos­o­phy Read­ing Group

Ahmed7 Jul 2026 4:55 UTC
3 points
0 comments1 min readLW link

The Geom­e­try of Yes: Map­ping Sy­co­phancy In­side an LLM’s Emo­tion Space

Pushpita Das7 Jul 2026 4:54 UTC
11 points
0 comments10 min readLW link

AI Safety Can’t Afford a Se­cond Cause

atlasaligned7 Jul 2026 4:48 UTC
37 points
9 comments3 min readLW link

Ranges of Prob­a­bil­ities: What Are They For?

SaltAndMetal7 Jul 2026 4:47 UTC
9 points
3 comments7 min readLW link

Ar­chi­tec­ture mat­ters for multi-agent security

7 Jul 2026 4:46 UTC
21 points
0 comments10 min readLW link

Data fil­ter­ing works a lot worse than you would ex­pect

7 Jul 2026 4:41 UTC
59 points
9 comments3 min readLW link

Fork Around and Find Out: Part 1—A Fork Snaps Into Ex­is­tence at Block 5

David Litman7 Jul 2026 4:36 UTC
11 points
0 comments7 min readLW link

growth tax

ofpetro6 Jul 2026 23:50 UTC
−9 points
0 comments17 min readLW link

Can we find whether mod­els have been back­doored?

6 Jul 2026 23:45 UTC
22 points
2 comments8 min readLW link

Info dump: Some Im­por­tant Models for Health and Fitness

benwr6 Jul 2026 23:40 UTC
98 points
14 comments22 min readLW link

Claude Code as a Claude Coach

Brendan Long6 Jul 2026 23:15 UTC
42 points
2 comments3 min readLW link
(www.brendanlong.com)

Be less open-minded: Against un­cal­ibrated open-mindedness

Nathan Young6 Jul 2026 23:00 UTC
25 points
2 comments2 min readLW link

A Re­view of An­thropic’s Global Workspace Paper

Neel Nanda6 Jul 2026 20:59 UTC
135 points
6 comments25 min readLW link

Vi­sion­ing: Con­cretely Imag­in­ing What You Want

6 Jul 2026 19:19 UTC
85 points
32 comments12 min readLW link

Per­son­hood for digi­tal minds is good

harsimony6 Jul 2026 18:48 UTC
9 points
0 comments8 min readLW link
(splittinginfinity.substack.com)

Train­ing AI to be bet­ter at cor­rect­ness than persuasion

MichaelDickens6 Jul 2026 18:07 UTC
24 points
1 comment2 min readLW link

A global workspace in lan­guage models

wesg6 Jul 2026 18:04 UTC
369 points
67 comments17 min readLW link
(www.anthropic.com)

Cur­rent views on large-scale longter­mist philanthropy

Zach Stein-Perlman6 Jul 2026 18:00 UTC
36 points
13 comments3 min readLW link

SFF is very suboptimal

Zach Stein-Perlman6 Jul 2026 18:00 UTC
118 points
10 comments5 min readLW link

Bound­ing eval aware­ness of ~hu­man-level AI across the safe-to-dan­ger­ous shift

6 Jul 2026 17:55 UTC
40 points
4 comments8 min readLW link

Sub-agent del­e­ga­tion chaining

David Rein6 Jul 2026 17:26 UTC
40 points
11 comments2 min readLW link

New fund­ing op­por­tu­nity on digi­tal minds from Longview

aog6 Jul 2026 16:30 UTC
14 points
0 comments2 min readLW link

Light­ning Jazz

Ben Goldhaber6 Jul 2026 16:01 UTC
21 points
0 comments3 min readLW link
(bengoldhaber.substack.com)

Tie train­ing can make DPO/​RLHF-trained AIs gen­er­al­ize better

6 Jul 2026 15:25 UTC
54 points
8 comments15 min readLW link
(arxiv.org)

Tips on lev­er­ag­ing AI for em­piri­cal research

Daniel Tan6 Jul 2026 13:50 UTC
30 points
0 comments7 min readLW link

Desider­ata for func­tional welfare ex­per­i­ments on LLMs

6 Jul 2026 12:34 UTC
34 points
1 comment15 min readLW link

Qualia-based emo­tion steer­ing makes llms at­tribute con­scious states to themselves

agastya6 Jul 2026 10:24 UTC
40 points
4 comments10 min readLW link

Com­pu­ta­tion en­ables Ac­tion: Ex­plod­ing the Si­mu­la­tion Fallacy

ChrisHibbert6 Jul 2026 2:54 UTC
27 points
1 comment6 min readLW link

Coun­ter­fac­tual mug­ging is a limit­ing case of Psy-kosh’s non-an­thropic problem

Abhimanyu Pallavi Sudhir6 Jul 2026 2:21 UTC
6 points
0 comments2 min readLW link

Get­ting an Oura ring im­proved my sleep and exercise

Daniel Tan5 Jul 2026 23:26 UTC
11 points
0 comments2 min readLW link

Fealty to Fidelity 👰

chaosmage5 Jul 2026 23:25 UTC
19 points
0 comments3 min readLW link

Reflec­tions on The Scout Mindset

James Brobin5 Jul 2026 20:18 UTC
12 points
0 comments3 min readLW link

When Gemma Thinks About Re­sources—it Fails: a Be­hav­ioral Experiment

TheVinci5 Jul 2026 19:23 UTC
8 points
2 comments2 min readLW link
(tarantulabs.com)

Claude’s mal­i­cious com­pli­ance and nor­mal­iza­tion of deviance

Steff5 Jul 2026 17:18 UTC
27 points
23 comments8 min readLW link

Book Re­view: The God Test

PeterMcCluskey5 Jul 2026 16:11 UTC
25 points
3 comments4 min readLW link

We need 3rd party Train­ing-Run Evaluations

Alex Meinke5 Jul 2026 15:55 UTC
176 points
2 comments11 min readLW link

Harry Pot­ter and the Rules of Quidditch

Tomás B.5 Jul 2026 14:32 UTC
161 points
8 comments3 min readLW link

A Nor­mal Ar­gu­ment for AI Risk

Silent Swift5 Jul 2026 9:32 UTC
16 points
3 comments8 min readLW link
(silentswift.substack.com)