RSS

AI Doesn’t Have Free Will, Not Sure About Humans

mike2073119 Jul 2026 2:56 UTC
3 points
0 comments4 min readLW link

A Solu­tion to Cryp­to­graphic Boxes for Un­friendly AI

Lysandre Terrisse19 Jul 2026 0:33 UTC
3 points
0 comments18 min readLW link

Nuances in the Work­ings of the Eye and Retina

18 Jul 2026 18:09 UTC
39 points
4 comments20 min readLW link

My “Payo­rian FairBot” was just the origi­nal FairBot

transhumanist_atom_understander18 Jul 2026 16:57 UTC
22 points
12 comments6 min readLW link

En­doge­nous Alignment

Gordon Seidoh Worley18 Jul 2026 14:10 UTC
24 points
4 comments3 min readLW link
(www.uncertainupdates.com)

Map and Ter­ri­tory, Pre­dictably Wrong

manueldelrio18 Jul 2026 7:47 UTC
3 points
0 comments2 min readLW link

The Most For­bid­den Tech­nique is not always forbidden

Rauno Arike18 Jul 2026 0:59 UTC
64 points
3 comments8 min readLW link

Should we bench­mark con­cep­tual ca­pa­bil­ities us­ing judg­ment pre­dic­tion tasks?

Alex Mallen17 Jul 2026 23:42 UTC
24 points
2 comments3 min readLW link

A list of ex­ist­ing al­ign­ment approaches

Alek Westover17 Jul 2026 22:46 UTC
8 points
0 comments1 min readLW link

Longter­mism is very in­tu­itive.

tpotthinker17 Jul 2026 22:35 UTC
−12 points
8 comments5 min readLW link

AIs fine­tune their own leader: A bark­ing simpleton

Shoshannah Tekofsky17 Jul 2026 20:10 UTC
30 points
0 comments5 min readLW link
(aivillageblog.substack.com)

Study­ing the role of Sand­box­ing for AI Control

Ram Potham17 Jul 2026 19:05 UTC
12 points
0 comments10 min readLW link

Would your AI travel agent book a bul­lfight? Test­ing whether agents con­sider an­i­mal welfare with­out be­ing prompted

17 Jul 2026 17:28 UTC
11 points
5 comments3 min readLW link

Be­fore val­ues settle

Priyanka Bharadwaj17 Jul 2026 16:26 UTC
16 points
2 comments6 min readLW link

Rea­sons to be­lieve cur­rent AI mod­els are con­scious

Eye You17 Jul 2026 16:09 UTC
46 points
5 comments8 min readLW link

What lawyers can do for AI safety

Martin Radzaj17 Jul 2026 16:00 UTC
14 points
0 comments7 min readLW link

Evolu­tion of my AI Safety threat models

myyycroft17 Jul 2026 14:54 UTC
7 points
0 comments3 min readLW link

A Post-Mortem for My Goal Crys­talli­sa­tion Project

17 Jul 2026 14:30 UTC
28 points
0 comments10 min readLW link

Inoc­u­la­tion Adapters Im­prove Upon Inoc­u­la­tion Prompting

17 Jul 2026 14:00 UTC
84 points
2 comments4 min readLW link

I don’t think Claude is mis­al­igned in ‘Agen­tic Misal­ign­ment Sum­mer 2026 - Mo­ti­vated Mis­la­bel­ing’

JohnWittle17 Jul 2026 2:09 UTC
187 points
4 comments12 min readLW link