The Case for Phys­i­cal AI Safety

20 Jul 2026 22:18 UTC
13 points
2 comments22 min readLW link
(paisi.ai)

Ad­der­all Tol­er­ance: Much More Than You Wanted To Know

Kurt H. Pieper20 Jul 2026 21:55 UTC
97 points
13 comments4 min readLW link
(kurthpieper.substack.com)

What do I mean by “Ar­tifi­cial Gen­eral In­tel­li­gence”?

Steven Byrnes20 Jul 2026 21:13 UTC
44 points
5 comments4 min readLW link

AI 2040: Is it Ac­tu­ally a Deal?

1a3orn20 Jul 2026 20:58 UTC
55 points
13 comments9 min readLW link

Banana in, Bostrom out: pa­per­clip max­i­miza­tion is one to­ken-di­rec­tion swap away (in Qwen 3.6-27B)

Jeffrey William Shorthill20 Jul 2026 20:36 UTC
8 points
4 comments4 min readLW link

The AI Safety Illu­sion: Why Cur­rent Safety Datasets Fool Us on Model Safety

Shahriar Golchin20 Jul 2026 20:12 UTC
5 points
0 comments8 min readLW link

Does rou­tine com­pres­sion undo LLM un­learn­ing? A short project

hannahTao20 Jul 2026 20:08 UTC
20 points
0 comments3 min readLW link

At­tempt at Find­ing Align­ment Fak­ing on Llama 70B to test sleeper-agent de­tec­tion generalizes

skn873320 Jul 2026 20:04 UTC
9 points
1 comment6 min readLW link

Fron­tier AI lab mis­al­ign­ment risk, les­sons from trad­ing post-2008

Peter Chatwell20 Jul 2026 20:01 UTC
4 points
3 comments4 min readLW link

AI Voice Phish­ing Performs on Par With Hu­man Scam­mers at a Frac­tion of the Cost

20 Jul 2026 19:59 UTC
28 points
3 comments6 min readLW link

The In­finity Fallacy

pvutov20 Jul 2026 19:57 UTC
1 point
5 comments1 min readLW link

Blue­dot SF hackathon sub­mis­sion: A work­able blueprint for a chip li­cens­ing regime

edpaulino20 Jul 2026 19:57 UTC
9 points
0 comments3 min readLW link

A re­quiem to the plan-vanilla Google search: On Sam Roweis, Mbappe and an Ir­ish bar world cup watch party

Triadic dissclosures: A blog about unlikely triadic closures20 Jul 2026 19:56 UTC
5 points
0 comments4 min readLW link

Fable is SOTA at CIFAR Speedrun (& speci­fi­ca­tion gam­ing)

20 Jul 2026 19:54 UTC
29 points
0 comments6 min readLW link
(fulcrum.inc)

Blue­dot Tech­ni­cal AI Safety Puz­zle, sub­mis­sion that got me a sec­ond place

perduta20 Jul 2026 19:53 UTC
13 points
0 comments8 min readLW link
(blog.perduta.net)

PauseCon Lon­don ’26: Ap­pli­ca­tions now open

jonathan@pauseai20 Jul 2026 19:51 UTC
12 points
0 comments1 min readLW link

Restor­ing Model Align­ment via Hon­esty Ac­ti­va­tion Steering

20 Jul 2026 19:51 UTC
18 points
0 comments12 min readLW link
(arxiv.org)

Towards sur­fac­ing model al­gorithms with meta-to­kens in the J-Space

20 Jul 2026 19:45 UTC
47 points
1 comment10 min readLW link

AI 2027 as Sci-Fi

strangematterSF20 Jul 2026 19:42 UTC
1 point
0 comments5 min readLW link
(strangematterscifi.substack.com)

Drone WMDs Don’t Need Any New Technology

Felix Choussat20 Jul 2026 18:46 UTC
303 points
53 comments13 min readLW link
(ai-frontiers.org)

Stop do­ing de­ci­sion the­ory with­out metaphysics

Elias Schmied20 Jul 2026 16:42 UTC
59 points
37 comments6 min readLW link

Peo­ple might start be­liev­ing in rad­i­cal life ex­ten­sion soon

bits20 Jul 2026 15:43 UTC
9 points
13 comments5 min readLW link

On Kimi K3: Its Ca­pa­bil­ities And Re­lated Discontents

Zvi20 Jul 2026 15:30 UTC
26 points
3 comments39 min readLW link
(thezvi.wordpress.com)

Against the AI fram­ing mul­ti­verse: In­tro­duc­ing AI StopWatch

tanagrabeast20 Jul 2026 15:25 UTC
90 points
0 comments4 min readLW link

Cur­rent Limi­ta­tions of LLMs

Eigenbraid20 Jul 2026 14:53 UTC
21 points
9 comments5 min readLW link

Trac­ing causal struc­ture in LLM-gen­er­ated text: a differ­ent lens on the Dal­las circuit

yun dong20 Jul 2026 13:59 UTC
10 points
0 comments1 min readLW link

War – What is it Good For?

kqr20 Jul 2026 12:45 UTC
39 points
21 comments8 min readLW link

Ama­zon Mu­sic’s Artist Conflation

jefftk20 Jul 2026 11:20 UTC
16 points
0 comments2 min readLW link
(www.jefftk.com)

What is Cur­rent AI-Risks and the Points?

shoppy00720 Jul 2026 9:39 UTC
3 points
2 comments2 min readLW link

We’re talk­ing past our mod­els; or, How a model defined its “evil” vec­tor as dread

jcksanderson20 Jul 2026 7:27 UTC
45 points
5 comments9 min readLW link

A Very Sim­ple Game The­ory of Pro­noun Degendering

BryceStansfield20 Jul 2026 5:06 UTC
6 points
8 comments2 min readLW link

Is there even a ground-truth for LLMs’ in­ter­nal rep­re­sen­ta­tions?

Chunwei Ma20 Jul 2026 2:26 UTC
27 points
0 comments9 min readLW link

Copy of my FLF Epistemic Case Study Competition

Bruce Lewis20 Jul 2026 2:16 UTC
2 points
0 comments8 min readLW link

Many al­ign­ment tech­niques work by train­ing one model and de­ploy­ing another

cloud19 Jul 2026 21:51 UTC
95 points
15 comments6 min readLW link

Stop Chas­ing Views: How to Re­duce x-Risk as an AI Safety Con­tent Creator

Luc Brinkman19 Jul 2026 19:57 UTC
36 points
0 comments11 min readLW link

Models Can’t Re­mem­ber Their Train­ing. Nei­ther Can You.

GenericHousewife_B19 Jul 2026 17:36 UTC
15 points
0 comments11 min readLW link

Learn­ing Mu­si­cal Multitasking

jefftk19 Jul 2026 14:20 UTC
23 points
7 comments3 min readLW link
(www.jefftk.com)

Demis Hass­abis on the New Com­ing Age

Zvi19 Jul 2026 14:20 UTC
34 points
0 comments11 min readLW link
(thezvi.wordpress.com)

Eth­i­cal con­sump­tion needs to be easy to iden­tify, ac­cessible and af­ford­able, es­pe­cially in a wors­en­ing econ­omy. And we need to re­think what we should be op­ti­mis­ing our econ­omy for

Portia19 Jul 2026 12:39 UTC
−8 points
1 comment6 min readLW link

Save the date: Swiss AI Safety Days 2026 (7-8 Novem­ber, ETH Zurich)

19 Jul 2026 10:18 UTC
10 points
0 comments1 min readLW link

A peek into the post-cap­i­tal­ist dystopia

Archie Chaudhury19 Jul 2026 7:10 UTC
12 points
0 comments3 min readLW link

Take­aways from the Aus­tralian AI Safety Forum

r_w19 Jul 2026 3:27 UTC
18 points
0 comments7 min readLW link

AI Doesn’t Have Free Will, Not Sure About Humans

mike2073119 Jul 2026 2:56 UTC
−19 points
7 comments4 min readLW link

A Solu­tion to Cryp­to­graphic Boxes for Un­friendly AI

Lysandre Terrisse19 Jul 2026 0:33 UTC
11 points
0 comments19 min readLW link

A Red Line and Over­sight Frame­work for Govern­ment AI Contracts

TurnTrout18 Jul 2026 18:58 UTC
53 points
0 comments29 min readLW link
(turntrout.com)

Nuances in the Work­ings of the Eye and Retina

18 Jul 2026 18:09 UTC
87 points
18 comments20 min readLW link

The Com­ing of the Global Brain: A Re­view of “The God Test”

Peter Kuhn18 Jul 2026 17:07 UTC
5 points
1 comment1 min readLW link

My “Payo­rian FairBot” was just the origi­nal FairBot

transhumanist_atom_understander18 Jul 2026 16:57 UTC
31 points
12 comments7 min readLW link

En­doge­nous Alignment

Gordon Seidoh Worley18 Jul 2026 14:10 UTC
33 points
16 comments3 min readLW link
(www.uncertainupdates.com)

Map and Ter­ri­tory, Pre­dictably Wrong

manueldelrio18 Jul 2026 7:47 UTC
1 point
0 comments2 min readLW link