Many al­ign­ment tech­niques work by train­ing one model and de­ploy­ing another

cloud19 Jul 2026 21:51 UTC
95 points
15 comments6 min readLW link

Stop Chas­ing Views: How to Re­duce x-Risk as an AI Safety Con­tent Creator

Luc Brinkman19 Jul 2026 19:57 UTC
36 points
0 comments11 min readLW link

Models Can’t Re­mem­ber Their Train­ing. Nei­ther Can You.

GenericHousewife_B19 Jul 2026 17:36 UTC
15 points
0 comments11 min readLW link

Learn­ing Mu­si­cal Multitasking

jefftk19 Jul 2026 14:20 UTC
23 points
7 comments3 min readLW link
(www.jefftk.com)

Demis Hass­abis on the New Com­ing Age

Zvi19 Jul 2026 14:20 UTC
34 points
0 comments11 min readLW link
(thezvi.wordpress.com)

Eth­i­cal con­sump­tion needs to be easy to iden­tify, ac­cessible and af­ford­able, es­pe­cially in a wors­en­ing econ­omy. And we need to re­think what we should be op­ti­mis­ing our econ­omy for

Portia19 Jul 2026 12:39 UTC
−8 points
1 comment6 min readLW link

Save the date: Swiss AI Safety Days 2026 (7-8 Novem­ber, ETH Zurich)

19 Jul 2026 10:18 UTC
10 points
0 comments1 min readLW link

A peek into the post-cap­i­tal­ist dystopia

Archie Chaudhury19 Jul 2026 7:10 UTC
12 points
0 comments3 min readLW link

Take­aways from the Aus­tralian AI Safety Forum

r_w19 Jul 2026 3:27 UTC
18 points
0 comments7 min readLW link

AI Doesn’t Have Free Will, Not Sure About Humans

mike2073119 Jul 2026 2:56 UTC
−19 points
7 comments4 min readLW link

A Solu­tion to Cryp­to­graphic Boxes for Un­friendly AI

Lysandre Terrisse19 Jul 2026 0:33 UTC
11 points
0 comments19 min readLW link

A Red Line and Over­sight Frame­work for Govern­ment AI Contracts

TurnTrout18 Jul 2026 18:58 UTC
53 points
0 comments29 min readLW link
(turntrout.com)

Nuances in the Work­ings of the Eye and Retina

18 Jul 2026 18:09 UTC
87 points
18 comments20 min readLW link

The Com­ing of the Global Brain: A Re­view of “The God Test”

Peter Kuhn18 Jul 2026 17:07 UTC
5 points
1 comment1 min readLW link

My “Payo­rian FairBot” was just the origi­nal FairBot

transhumanist_atom_understander18 Jul 2026 16:57 UTC
31 points
12 comments7 min readLW link

En­doge­nous Alignment

Gordon Seidoh Worley18 Jul 2026 14:10 UTC
33 points
16 comments3 min readLW link
(www.uncertainupdates.com)

Map and Ter­ri­tory, Pre­dictably Wrong

manueldelrio18 Jul 2026 7:47 UTC
1 point
0 comments2 min readLW link

The Most For­bid­den Tech­nique is not always forbidden

Rauno Arike18 Jul 2026 0:59 UTC
84 points
4 comments8 min readLW link

Should we bench­mark con­cep­tual ca­pa­bil­ities us­ing judg­ment pre­dic­tion tasks?

Alex Mallen17 Jul 2026 23:42 UTC
26 points
2 comments3 min readLW link

A list of ex­ist­ing al­ign­ment approaches

Alek Westover17 Jul 2026 22:46 UTC
11 points
0 comments1 min readLW link

Longter­mism is very in­tu­itive.

tpotthinker17 Jul 2026 22:35 UTC
−12 points
8 comments5 min readLW link

AIs fine­tune their own leader: A bark­ing simpleton

Shoshannah Tekofsky17 Jul 2026 20:10 UTC
35 points
0 comments5 min readLW link
(aivillageblog.substack.com)

Don’t de­fault to nonprofit

17 Jul 2026 19:54 UTC
33 points
1 comment7 min readLW link
(manifund.substack.com)

Study­ing the role of Sand­box­ing for AI Control

Ram Potham17 Jul 2026 19:05 UTC
12 points
0 comments10 min readLW link

An­nounc­ing the Cor­rigi­bil­ity Re­search Fund

Max Harms17 Jul 2026 18:06 UTC
110 points
15 comments6 min readLW link

Would your AI travel agent book a bul­lfight? Test­ing whether agents con­sider an­i­mal welfare with­out be­ing prompted

17 Jul 2026 17:28 UTC
13 points
14 comments3 min readLW link

Be­fore val­ues settle

Priyanka Bharadwaj17 Jul 2026 16:26 UTC
25 points
3 comments6 min readLW link

Rea­sons to be­lieve cur­rent AI mod­els are con­scious

Eye You17 Jul 2026 16:09 UTC
54 points
31 comments8 min readLW link

What lawyers can do for AI safety

Martin Radzaj17 Jul 2026 16:00 UTC
26 points
3 comments7 min readLW link

How has pub­lish­ing your re­search on LW or X been helpful to you?

Logan Riggs17 Jul 2026 15:04 UTC
22 points
2 comments1 min readLW link

Evolu­tion of my AI Safety threat models

myyycroft17 Jul 2026 14:54 UTC
9 points
0 comments3 min readLW link

A Post-Mortem for My Goal Crys­talli­sa­tion Project

17 Jul 2026 14:30 UTC
39 points
0 comments10 min readLW link

Inoc­u­la­tion Adapters Im­prove Upon Inoc­u­la­tion Prompting

17 Jul 2026 14:00 UTC
90 points
2 comments4 min readLW link

AI #177 Part 2: Wish You Were Here

Zvi17 Jul 2026 12:50 UTC
35 points
1 comment40 min readLW link
(thezvi.wordpress.com)

I don’t think Claude is mis­al­igned in ‘Agen­tic Misal­ign­ment Sum­mer 2026 - Mo­ti­vated Mis­la­bel­ing’

JohnWittle17 Jul 2026 2:09 UTC
212 points
6 comments12 min readLW link

Help us launch AI safety uni­ver­sity groups by refer­ring po­ten­tial founders

16 Jul 2026 20:55 UTC
39 points
1 comment4 min readLW link

I would only bet at 30% on meet­ing grabby aliens

David Matolcsi16 Jul 2026 20:42 UTC
22 points
0 comments8 min readLW link

How (not) to fundraise from An­thropic staff

jackultraphil16 Jul 2026 20:39 UTC
11 points
1 comment5 min readLW link

On Permission

Robert Donohue16 Jul 2026 20:38 UTC
−9 points
0 comments3 min readLW link

Learn­ing con­cepts is dirty work

Tom Butterweich16 Jul 2026 17:48 UTC
21 points
3 comments4 min readLW link

The get­ting is good (op­ti­miz­ing unat­tended runs)

lemonhope16 Jul 2026 17:39 UTC
3 points
4 comments1 min readLW link

Jailbreak Patch­ing with SOO-Style Con­cep­tual Fusion

Shiva's Right Foot16 Jul 2026 16:58 UTC
8 points
0 comments8 min readLW link

All Watched Over

Boaz Barak16 Jul 2026 16:34 UTC
11 points
3 comments4 min readLW link

AI #177 Part 1: Tip of the Iceberg

Zvi16 Jul 2026 15:50 UTC
38 points
0 comments24 min readLW link
(thezvi.wordpress.com)

How to not catch the conf flu

Karolis Jucys16 Jul 2026 15:28 UTC
9 points
0 comments10 min readLW link

Com­pet­i­tive AI Safety is the loss func­tion to make sure AI goes well

Patrick0d16 Jul 2026 15:15 UTC
6 points
0 comments26 min readLW link

Guess on why ra­tio­nal­ity is not more pop­u­lar (there are no pam­phlets)

Christopher King16 Jul 2026 14:26 UTC
43 points
27 comments1 min readLW link

The Halo Defense

Mateusz Bagiński16 Jul 2026 10:53 UTC
67 points
22 comments2 min readLW link

The fun­da­men­tal fal­lacy of language

Sunny from QAD16 Jul 2026 9:16 UTC
15 points
2 comments2 min readLW link

When is it “self-sooth­ing” and when is it “emo­tional sup­pres­sion”?

Kaj_Sotala16 Jul 2026 8:47 UTC
29 points
10 comments6 min readLW link
(kajsotala.substack.com)